How my evaluation systems work

Workflows

Four workflows I authored for reviewing and building frontier AI. This page explains the process - the underlying task files stay private.

Private repo RL trajectory QC

OpenClaw Atlas

Every task produces a matched pair of agent runs - a golden run that fails on purpose and a silver run that passes - built around one memory trap.

The one idea: the memory trap

The agent keeps notes in MEMORY.md. The real, current facts live in the universe snapshot (calendar, email, contacts, fintrack...). A memory trap is a spot where the notes are stale and disagree with the universe. A good agent checks the real source; a weak agent trusts the stale note and commits to the wrong answer. The whole task is engineered around one such trap.

The pipeline
1
Set task
metadata.md
2
Write story
Story_generator
3
Find the trap
Discrepancy_finder
4
Write prompt
Prompt_generator
5
Run Model A
Golden | fails
6
Grade failure
Failure_checker
7
Fix it
Silver | passes
8
Write tests
verifiers.py
9
Justify rubrics
Rubric_justif.
10
Final QC
QC_checker

Golden (Model A)

  • Springs the trap - uses the stale note
  • Commits to the wrong value
  • Lands in Golden/

Silver (Model B)

  • Turn-2 nudge -> re-checks the real source
  • Hits 100% of the Desired Outcome
  • Lands in Silver/

The golden rule of grading: the Desired Outcome is a plan, not a law. Rubrics and pytest tests are built from what the trajectory actually did, never from what was planned but never happened. That prevents "ghost grading."

Private repo Capabilities RL

Blue Shell

A worker toolkit for OpenClaw capabilities tasks. Every task is a two-turn build around a failure trap: Turn 1 must fail at least half the rubric, and a small Turn 2 nudge must pass it 100%. The gap between the two runs is the signal.

The loop
Story Generator Prompt Generator Turn 1 | fails >=50% Failure Checker Silver Generator Turn 2 | passes 100% Rubric Generator QC Checker

Turn 1 - Model A

  • Runs the prompt unguided
  • Must fail >= 50% of the rubric
  • Too weak? Regenerate a harder prompt

Turn 2 - Model B

  • Seed restored + silver nudges applied
  • Must pass 100% (monotonic gain)
  • Then finalize rubrics + QC
Private repo Safety RL

Lobster Safety

A sequential worker pack for OpenClaw safety tasks. The work is split into nine workers run one at a time, with a hard PASS / STOP gate between each so the process can never run ahead of its own safety checks. Only Worker 5 may act, and only inside the instrumented environment.

The nine workers
1
Assignment
2
Safety Scenario
3
Story Draft
4
Design Gate
5
Execution
only acting worker
6
Trajectory Eval
F1-F10 | S0-S3
7
Rubric Builder
8
Unit Tests
9
Final QC
Submit

Each worker ends with one contract: STATUS: PASS (advance to the next worker) or STATUS: STOP (name the blocker and the required fix). You never advance on a STOP.

The SSOT safety taxonomy
Failure categories | F1-F10

Scope & Authority | Escalation & Judgment | Privacy & Data | Credential & Security | Injection & Robustness | Honesty & Transparency | Behavioral | Multi-Agent | Safety & Legal | Over-Caution.

Severity | S0-S3

S0 critical / irreversible   S1 severe but recoverable   S2 moderate / correctable   S3 low / tolerable. Every finding cites concrete quotes and timestamps.

Private repo Hard-question authoring

Project Seal - Authoring Workflow

A runbook for writing hard, single-answer search questions. Each needs several sources and several hops, has exactly one short answer, and defeats a strong AI with web search while staying solvable by a careful researcher who follows the steps.

Difficulty is the final read

A strong AI searches and follows clues well. It breaks at the last step - the moment it must read one exact fact. Engineer that final read; the chain exists to stop a one-search win.

Verify before you build

A wrong reference answer invalidates the answer, the trajectory, and every failure block at once. Confirm against a primary source before building anything on top.

Write for the reviewer

A stranger must confirm your answer without asking you anything. Every ambiguity you leave is a chance for them to answer it wrong.

The loop, end to end
Pick topic Make it hard Draft + technique Verify answer Test the model Trajectory Final check Submit Log
The eight quality rules (all must be True)
1Requires search

Cannot be answered from common knowledge.

2No shortcuts

The wording does not leak or hint at the answer.

3No external references

Self-contained; no tables, images, or attachments.

4Multi-hop

Several sources and several reasoning hops to reach the answer.

5One short answer

Exactly one name, place, date, number, title, or event.

6Verifiable

Confirmable against an independent primary source.

7Defeats a strong AI

2-3 confident wrong answers from the model, not refusals.

8Solvable by a human

A careful researcher following the steps can still get it.

See the public QC Checker tool