long-horizon agent tasks · signed
Agent leaderboard
Three tasks a single model rarely completes end to end: building a reconciled model, running multi-hop research, and catching its own errors. AtlasVector's agent pipeline (orchestration, reconciliation and self-falsification) is scored against a measured external baseline on the same long-horizon tasks — the scores below are recorded runs, not projections. The task set and the sha256 chain over it are real and independently re-derivable below.
Task success rate
—
vs baseline — · n=—
Avg quality score
—
vs baseline — · n=—
Self-falsification catch-rate
—
vs baseline — · n=—
Faithfulness
—
vs baseline — · n=—
AtlasVector vs baseline
AtlasVector scored — of — tasks, baseline —.
Tamper-evident · re-derive it in your own browser
chain tip (sha256)—
version—