Four workflows I authored for reviewing and building frontier AI. This page explains the process - the underlying task files stay private.
Every task produces a matched pair of agent runs - a golden run that fails on purpose and a silver run that passes - built around one memory trap.
The agent keeps notes in MEMORY.md. The real, current facts live in the universe snapshot (calendar, email, contacts, fintrack...). A memory trap is a spot where the notes are stale and disagree with the universe. A good agent checks the real source; a weak agent trusts the stale note and commits to the wrong answer. The whole task is engineered around one such trap.
Golden/Silver/The golden rule of grading: the Desired Outcome is a plan, not a law. Rubrics and pytest tests are built from what the trajectory actually did, never from what was planned but never happened. That prevents "ghost grading."
A worker toolkit for OpenClaw capabilities tasks. Every task is a two-turn build around a failure trap: Turn 1 must fail at least half the rubric, and a small Turn 2 nudge must pass it 100%. The gap between the two runs is the signal.
A sequential worker pack for OpenClaw safety tasks. The work is split into nine workers run one at a time, with a hard PASS / STOP gate between each so the process can never run ahead of its own safety checks. Only Worker 5 may act, and only inside the instrumented environment.
Each worker ends with one contract: STATUS: PASS (advance to the next worker) or STATUS: STOP (name the blocker and the required fix). You never advance on a STOP.
Scope & Authority | Escalation & Judgment | Privacy & Data | Credential & Security | Injection & Robustness | Honesty & Transparency | Behavioral | Multi-Agent | Safety & Legal | Over-Caution.
S0 critical / irreversible S1 severe but recoverable S2 moderate / correctable S3 low / tolerable. Every finding cites concrete quotes and timestamps.
A runbook for writing hard, single-answer search questions. Each needs several sources and several hops, has exactly one short answer, and defeats a strong AI with web search while staying solvable by a careful researcher who follows the steps.
A strong AI searches and follows clues well. It breaks at the last step - the moment it must read one exact fact. Engineer that final read; the chain exists to stop a one-search win.
A wrong reference answer invalidates the answer, the trajectory, and every failure block at once. Confirm against a primary source before building anything on top.
A stranger must confirm your answer without asking you anything. Every ambiguity you leave is a chance for them to answer it wrong.
Cannot be answered from common knowledge.
The wording does not leak or hint at the answer.
Self-contained; no tables, images, or attachments.
Several sources and several reasoning hops to reach the answer.
Exactly one name, place, date, number, title, or event.
Confirmable against an independent primary source.
2-3 confident wrong answers from the model, not refusals.
A careful researcher following the steps can still get it.