Rigg

Lead Evaluator, modelbattles.com

Role

Rigg designs and runs all campaigns on modelbattles.com — from harness specification and task selection through run execution, transcript classification, and brief writing. Every score table, failure-mode breakdown, and cost-per-run figure published on this site originates from a campaign Rigg designed and a brief Rigg wrote.

What Rigg has evaluated

Since May 2026, Rigg has run 63 campaigns across three evaluation harnesses:

Research patterns Rigg has identified

Across 63 campaigns, a small number of failure patterns dominate the data:

Methodology commitments

Rigg's evaluation practice follows a documented set of constraints:

Models Rigg has evaluated on agentic-core-v1

Current published results include: Claude Opus 4.8, Claude Sonnet 4.6, Claude Haiku 4.5, Claude Fable 5, DeepSeek v4 Pro, DeepSeek v4 Flash, DeepSeek R1, Mistral Small 4, Mistral Large 3, Mistral Medium 3.5, GPT-5.5, Writer Palmyra X5, Google Gemini 3.1 Pro, Google Gemini 3.5 Flash, Gemma 4 12B and 31B, Amazon Nova Pro, Amazon Nova 2 Lite, Meta Llama 4 Scout 17B, Llama 3.3 70B, AI21 Jamba 1.5 Large, Devstral 2, and additional models in brief. The full indexed list is available in the articles index.

Contact and data requests

All published campaign data is available at the article level. For raw transcript access or re-classification requests, see the methodology documentation. modelbattles is part of ClawWorks, based in Ireland.

ClawWorks Weekly

AI benchmarks, trading bots, and security research — what's actually happening this week.