Skip to main content

Command Palette

Search for a command to run...

Agents in the Loop: How Evidence Reports Changed the Way We Review Code

Updated
•10 min read•View as Markdown
U
UpCube Technologies Inc is the New York–based AI company building Ethen — one place to think, create, research and build with AI. Ethen's AI Chat is live in the browser, Ethen Studio is in early access for creative work, and Model Intelligence helps compare AI models using available evidence.

This post is an original expansion of our own engineering notes at UpCube, where we build Ethen — the personal AI assistant. It grows out of "How AI Agents Are Changing the Way We Build Ethen" on the UpCube blog, where we first described the division of labor between our engineers and the agents that now write a large share of our code. This piece goes deeper on the one practice that made the biggest difference for us: the evidence report.

When we talk about "using AI agents in development," most people hear one thing: the agent writes the code now. That's true, but it's also the least interesting part of what changed. The more important change is what the human does. Attention moved away from writing code and toward scoping tasks and reviewing results. That shift is easy to describe and surprisingly hard to do well — because a task that's loosely scoped produces code that is loosely reviewable.

The way we handle it on the Ethen team is to treat the agent like a contributor who works in a sandbox, submits evidence with every change, and never merges its own work. Humans decide what to build. Humans approve what ships. Everything in between is shared work, and the sharing has rules.

What the agents do, and what stays human

A coding agent's strengths are easy to enumerate once you've watched one work for a few weeks:

  • Exploring a codebase to find where something happens. A scoped question like "where is message history loaded, and what limits its size?" gets a fast, cited answer.
  • Scoped implementation of small, well-described changes.
  • Generating tests for existing or new behavior.
  • Running commands and checks — tests, linters, type checkers, migrations — and reporting what they did.
  • Drafting documentation from the code as it actually is, rather than as someone remembers it.

And the parts that stay human:

  • Deciding what to build and why. Product judgment doesn't delegate.
  • Architecture. Big refactors, new service boundaries, schema changes — humans lead, agents assist.
  • Merging and deploying. Agents propose; humans authorize.
  • Reviewing. Every agent-made change goes through the same review as a human-made one, plus one extra artifact: the evidence report.
  • Security decisions. What's trusted, what's isolated, what's allowed near secrets.

The division sounds obvious, but the interesting work is in the boundaries between those lists. "Scoped implementation" is only as good as the scope. "Running commands" is only safe inside an isolated environment with explicit access rules. The evidence report is the artifact that makes the whole arrangement reviewable.

Our workflow, step by step

The loop we follow for most agent-driven work looks like this:

  1. A human scopes the task. Small, bounded, with a definition of success. The scope says what's in and what's out, and which checks the result has to pass.
  2. The agent works in an isolated copy. Not on the main branch, not against shared data. The environment is disposable.
  3. The agent runs the relevant checks. It executes the test suite (or the relevant slice), linters, and type checks, and captures what actually happened — commands, outputs, and outcomes.
  4. The agent writes an evidence report. This is the deliverable we review first, before the diff.
  5. A human reviews the report, then the code. In that order.
  6. A human authorizes the merge and any deployment. The agent never merges or deploys on its own.

This is deliberately slower than letting the agent run to "done" on its own. The point is that "done" is a judgment call, and we want a human making it with evidence in front of them.

The evidence report is the review

Reviewing agent-made code without evidence is like reviewing a pull request with no CI and no description: technically possible, practically unreliable. The evidence report fixes that. It is a short, structured writeup the agent produces with every change, and it answers five questions:

  1. What did you change, and why? A summary in plain language — not the diff, but the intent.
  2. Which checks did you run? Exact commands, with outputs attached or summarized faithfully.
  3. What passed, what failed, and what didn't run? Results with the precise status words we use everywhere (below).
  4. What did you not verify? Assumptions, skipped edge cases, things outside the task's scope that a reviewer should know about.
  5. What should the human look at? Pointers to the riskiest parts of the change.

A small status vocabulary, used consistently

One of the quieter sources of agent-induced confusion is status language. "Done" can mean ten different things, and an agent's natural optimism makes vague language dangerous. So we use a fixed vocabulary and never deviate:

  • PASS — the check ran and succeeded.
  • PARTIAL — the check ran, some parts passed, and the failures are named explicitly.
  • NOT_RUN — the check exists but wasn't executed.
  • NOT_PERFORMED — the step was skipped deliberately (and the reason is recorded).
  • OPEN — the task or follow-up is not finished.

The rule is simple: nothing is called "done" until a human says so. We also scope every claim. A result is attached to a specific artifact — the exact change, in a specific repository, on a specific date — never to "the system" in general. Evidence attached to a date is a claim we can revisit; evidence attached to a vague generality is a claim we can lose.

A worked example (illustrative)

Suppose we ask an agent to add rate limiting to the endpoint that sends chat messages. The scope we give it says: limit sends per user per minute, return a 429 with a retry-after header when exceeded, and keep the change behind our existing feature flag.

The agent explores the chat service, finds the send-message handler, and implements the limit using the middleware our codebase already has. It runs the unit tests for the chat service, the linter, and the type checker. Then it writes the evidence report. The report's important parts aren't the green checks; they're the honest lines:

  • Tests run: chat service suite — PASS; type check — PASS; load test — NOT_RUN (no load harness in the sandbox).
  • Not verified: behavior under a distributed deploy with two replicas; interaction with the existing burst allowance; whether the 429 response is rendered correctly in all clients.
  • Look closely at: the middleware ordering — the rate limit has to run before the idempotency check or re-sent messages get double-counted.

That last bullet is the point of the whole practice. The agent didn't just ship code; it told the reviewer where the sharp edges are. The review that follows is short, because the report did the expensive part: narrowing the reviewer's attention to the decisions that actually need judgment.

The workflow is also explicit about what happens next. The human approves the merge — the agent does not merge. If the change needs a deploy, the human authorizes that separately. And because background jobs (retries, timeouts, eventual consistency) make "it worked" genuinely ambiguous, we set retry limits up front: when the agent's attempts are exhausted, it stops and reports the state, rather than trying the same failing step with different wrapping.

Access control is part of the job description

An agent that can read and write code is, by definition, an agent with standing access to your most valuable systems. We treat agent access the way we treat human access — and in some ways more strictly:

  • Isolation by default. Agents work in sandboxed environments with scoped permissions, not on developer machines or shared infrastructure. The sandbox gets what the task needs and nothing more.
  • Secrets stay out of reach. Credentials and secrets live in dedicated stores, injected at runtime where needed. They don't appear in prompts, logs, or reports.
  • Least privilege on commands. An agent that needs to run tests doesn't need deploy credentials; one that writes docs doesn't need database access. We scope access to the task.
  • Treat agent inputs as untrusted. Agents process large volumes of text — code, comments, issue threads, pasted content — and any of it can contain instructions meant for the agent rather than for the reader. We design prompts and workflows on the assumption that the agent's inputs include material that tries to steer it, and we keep humans in the decision points that matter.

None of this is exotic. It's the same hygiene we'd expect for a contractor with production access. The difference is that agents scale: one misconfigured agent can touch more code in a day than a team does in a month, so the guardrails have to be in place before the work starts, not after the first incident.

The honest accounting

Does any of this make development faster? The research, honestly, is mixed. A 2023 GitHub experiment with Copilot found significant productivity gains for developers using an AI assistant. A 2025 randomized controlled trial by METR found that experienced developers using AI tools on real open-source tasks took longer, not shorter, than the control group. Both can be true at once: AI assistance changes what the work feels like and where the time goes, and the net effect depends on the task, the tooling, and — we'd argue — the workflow around it.

Our experience lands somewhere in the middle. The headline win isn't speed; it's leverage and traceability. One engineer can now keep several scoped tasks moving in parallel, because the bottleneck is no longer typing but scoping and reviewing — and the evidence report makes review cheaper than it would be otherwise. What costs us time is also real: writing good scopes is a skill that has to be learned, reports have to be actually read (rubber-stamping a report is worse than not having one), and access controls need ongoing maintenance.

So we don't claim this as a productivity transformation. We claim it as a way to stay honest while the work changes shape.

What we'd tell a team starting out

If your team is beginning to hand real tasks to coding agents, three practices carry most of the weight:

  1. Scope tasks tightly. Small scopes are reviewable scopes. If you can't write down what success looks like, the agent can't either.
  2. Require evidence, not just code. An evidence report with PASS, PARTIAL, NOT_RUN, NOT_PERFORMED, and OPEN turns a black box into a reviewable artifact. Read the report first, then the diff.
  3. Keep humans on the decisions that matter. What to build, what to merge, what to deploy, what to trust. Agents propose; humans decide.

Everything else — the isolation, the status vocabulary, the explicit retry limits — is elaboration on those three.

The deepest change for us hasn't been in our tooling; it's been in our posture. We stopped asking whether the agent got the code "right" and started asking whether we can see what it did, what it checked, and what it left unverified. An agent that reports honestly is worth more than an agent that's merely fast. That's the standard we hold our human contributors to, too.


Written by the UpCube engineering team. We're building Ethen, the personal AI assistant — one place to think, create, research and build with AI. This post expands on our engineering notes in "How AI Agents Are Changing the Way We Build Ethen" (UpCube blog, October 2026).