AI coding-agent reliability

Know if the agent really finished.

Paste Codex, Claude Code, Gemini CLI, OpenCode, or Qwen logs and get a run receipt with false-success, loop, context, cost, and validation warnings in under a minute.

01

False-success scan

Flags "should work", skipped validation, unverified fixes, and other language that makes agent output risky.

02

Failure radar

Finds command errors, test failures, destructive operations, dirty worktrees, and repeated retry loops.

03

Context and spend

Estimates token pressure and spend from logs so long sessions do not hide runaway context or cost.

04

Shareable receipt

Exports a markdown receipt for PRs, incident notes, client handoffs, and your own sanity checks.

Workflow

Three steps from messy log to decision.

  1. Paste or uploadDrop in a session transcript, terminal log, or JSONL text export.
  2. Analyze riskAgentFlight scores the run and groups evidence into review categories.
  3. Export receiptSave the markdown receipt and run the recommended final checks.

Try it now

Analyze an AI-agent run

Status Waiting
Score --
Findings 0

Paste a log to generate a reliability receipt.


          

Pricing

Start free, pay when receipts become workflow.

Free

$0

  • Single-log browser analysis
  • Markdown receipt export
  • No account required

Pro

$19/mo

  • Saved projects and history
  • GitHub PR receipt comments
  • Provider-specific parsers
  • Team policy rules