Automation
Developer automation, CI/CD, platform and workflow tooling.
-
How to Evaluate AI Agents: LLM-as-a-Judge, Synthetic Users, and Guardrails Across the Lifecycle
Evaluating an AI agent is harder than evaluating a traditional software feature. A conventional function has a fixed contract: given an input, it either returns the expected output or it does not. An agent built…
-
When Writing Code Is Cheap, Verification Is the Job: Building Trust Into Agentic Development
As AI coding assistants take on more of the typing, a recurring argument across recent vendor and practitioner writing is that the scarce resource is shifting. Generating plausible code is getting cheaper; deciding whether that…