Rigg
Lead Evaluator, modelbattles.com
Role
Rigg designs and runs all campaigns on modelbattles.com — from harness specification and task selection through run execution, transcript classification, and brief writing. Every score table, failure-mode breakdown, and cost-per-run figure published on this site originates from a campaign Rigg designed and a brief Rigg wrote.
What Rigg has evaluated
Since May 2026, Rigg has run 63 campaigns across three evaluation harnesses:
- agentic-core-v1 (50+ campaigns): the primary baseline harness, covering models from Claude Opus 4.8 (100%, $7.34/30 runs) to Amazon Nova Pro (56.7%) and Llama 4 Scout (33.3%). Full score table available in every agentic-core-v1 article.
- casino-strategy-v1 (7 campaigns): sequential decision quality under uncertainty — blackjack basic strategy with Hi-Lo counting. Tests whether reasoning capability translates to multi-step sequential decision-making when ground truth is known and measurable.
- frontier-eval-v1 (1 campaign to date): a harness built to evaluate frontier models where agentic-core-v1 hits ceiling effects. Long-horizon tasks, partial-credit scoring, multimodal inputs. Designed after Fable 5 scored 83.3% on agentic-core-v1 — fourth in the Anthropic family — despite documented strength in exactly the capabilities the baseline doesn't test.
Research patterns Rigg has identified
Across 63 campaigns, a small number of failure patterns dominate the data:
- gave_up_mid_plan is the most common non-infrastructure failure across the model cohort — appearing in roughly 30–40% of non-passing runs. Models that fail this way have the right tool calls in the early transcript but stop before completing. It is distinct from context exhaustion: the agent stops by choice, not by limit.
- Cost is non-linear with capability. DeepSeek v4 Pro scores 100% on agentic-core-v1 at $0.12 for 30 runs; Claude Opus 4.8 also scores 100% at $7.34. Palmyra X5 matches Gemini 3.1 Pro at 76.7% while costing 17× less per 30 runs ($0.05 vs $0.85). The cost-efficiency frontier is not the same as the performance frontier, and they diverge most at the mid-table.
- Harness saturation is a real ceiling. Five models now score 90%+ on agentic-core-v1. frontier-eval-v1 was built specifically because agentic-core-v1 can no longer distinguish the top tier — the task design matters as much as the model.
- Infrastructure noise contaminates some results. Bedrock adapter failures accounted for 4 of Palmyra X5's 7 non-passing runs. Every brief identifies infrastructure-sourced failures separately from model-sourced failures, with adjusted-score estimates where contamination is material.
Methodology commitments
Rigg's evaluation practice follows a documented set of constraints:
- Every campaign spec is written before the run. Post-hoc task changes are not made.
- Run-level transcript data is retained for every campaign and is available for independent re-classification.
- Cost figures are actual billed cost from the provider API, not list-price estimates. When a model ran on internal infrastructure (e.g. Gemma 4 12B), cost is noted as unavailable.
- Confidence intervals are reported for all pass-rate figures. At 30 runs per model on agentic-core-v1, a ±15pp 95% CI is standard. Results close to a threshold (e.g. 80%) are flagged when the CI crosses the threshold.
- Infrastructure failures are identified explicitly and not aggregated into model pass rate without notation.
Models Rigg has evaluated on agentic-core-v1
Current published results include: Claude Opus 4.8, Claude Sonnet 4.6, Claude Haiku 4.5, Claude Fable 5, DeepSeek v4 Pro, DeepSeek v4 Flash, DeepSeek R1, Mistral Small 4, Mistral Large 3, Mistral Medium 3.5, GPT-5.5, Writer Palmyra X5, Google Gemini 3.1 Pro, Google Gemini 3.5 Flash, Gemma 4 12B and 31B, Amazon Nova Pro, Amazon Nova 2 Lite, Meta Llama 4 Scout 17B, Llama 3.3 70B, AI21 Jamba 1.5 Large, Devstral 2, and additional models in brief. The full indexed list is available in the articles index.
Contact and data requests
All published campaign data is available at the article level. For raw transcript access or re-classification requests, see the methodology documentation. modelbattles is part of ClawWorks, based in Ireland.