CryptoPatrick / code

Code · AI evaluation

Ripley

A daemon that keeps checking whether an LLM coding agent is still trying.

What it is

Model quality can change from one week to the next without notice. Ripley is a small daemon that runs simple, deterministic test tasks against the Claude Code CLI on a schedule, scores each run on correctness, token use and duration, and stores the results in SQLite. Rolling statistics flag when the pass rate drops. It is named after Ellen Ripley, who follows procedure and notices when systems that claim to be fine aren’t.

Every benchmark has a token budget (200 tokens by default), so monitoring doesn’t cost much.

Using it

# config.yaml
daemon:
  interval: "30m"            # how often to run benchmarks
  db_path: "./ripley.db"
claude:
  model: "Sonnet"
  default_max_tokens: 200
monitoring:
  rolling_window: 10         # runs per rolling statistic
  warning_threshold: 0.7     # warn if the pass rate falls below 70%
make build
./ripleyd        # run the daemon
./ripleyctl      # inspect results and warnings

Ideas for using it

  • Regression alarm for your own prompts: replace the sample tasks with the prompts your product depends on.
  • Compare models or settings over a week instead of on one lucky run.
  • Evidence for a vendor conversation: a log of pass rates over time beats an impression.

Status

Prototype; works with the Claude Code CLI only.