Retrieved article excerpt
Open article · Retrieved 2026-09-11T09:23:06.449895+00:00
Artificial Analysis K Optima Build your own custom benchmark Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency Try Optima How it works Contract Review Benchmark Build agent drafting tasks Benchmark how well models review our supplier contracts — flag risks, extract terms. Working Reading your example contracts Drafting grading rubrics Writing task 18 of 24 Contract Review Benchmark 24 tasks · rubrics attached · 4 categories 1 . Build 2 . Run 3 . Grade Why build your benchmark with Optima Results specific to your use case Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released. Cut cost and time by over 10x Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models Built on Artificial Analysis grading expertise Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head How it works 1 Give Optima your context Describe it Import your traces Use your coding agent Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you. Describe your use case | Attach examples pptx AA-Merch_Sales_Pitch_Deck pptx AA-Merch_Startup_Swag_Pitch two decks we were happy with — this is what good looks like Build agent drafts tasks + rubrics Your drafted benchmark Sales Deck Creation 5 tasks · 10 criteria · head-to-head Northstar Hotels welcome-kit pilot Crescent Arts Museum collection pitch Apex Trails 12-month merch programme Tidepool summer pre-season launch City Sound Festival activation 2 Choose your evaluation type Task style Objective Rubric judge Subjective Optima Q&A The model answers questions that have known correct answers Document input The model answers questions about your uploaded files Agentic The model completes tasks and produces deliverables, using tools along the way Interaction Simulate a real conversation with different types of users Coming soon Q&A Objective grading Example task prompt Which HS tariff code applies to lithium-ion e-bike batteries? How it's graded The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion. 3 Run across models Your benchmark 24 tasks Flag the risky clauses Extract the payment terms Draft the counterparty summary same tasks, same conditions, every model Run sandbox + tools Models you picked Claude Fable 5 GPT-5.6 Sol Kimi K3 Gemini 3.6 Flash tokens, cost and time recorded per task 24/24 Or bring your own agent Your own agent can compete in the same run, over HTTP. Claude Opus 4.6 GPT-5.2 Gemini 3 Pro your-agent POST /runs/9f2c/artifacts Benchmark run 24 tasks · same judges Leaderboard 1 Claude Opus 4.6 3 GPT-5.2 4 Gemini 3 Pro 2 your-agent 4 Grade and decide Contract Review Benchmark 24 tasks · 4 models Model Score Cost / Task Time / Task Claude Fable 5 60 $ 0.18 74 s GPT-5.6 Sol 59 $ 0.11 52 s Kimi K3 57 $ 0.04 61 s Gemini 3.6 Flash 50 $ 0.02 29 s strong and cheap Score Cost per task $ 0.00 $ 0.05 $ 0.10 $ 0.15 $ 0.20 Claude Fable 5 60 · 74 s per task GPT-5.6 Sol 59 · 52 s per task Kimi K3 57 · 61 s per task Gemini 3.6 Flash 50 · 29 s per task Pricing Building and running a benchmark is priced on token usage. Grading is priced per unit judged. Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top. Rubric grading is $0.002 per criterion, per model with standard judges, or $0.040 with premium judges. Pairwise grading is $0.006 per match with standard judges, or $0.150 with premium judges. Rubric grading uses one judge from the selected tier; pairwise grading uses the full judge panel. Standard uses high-capability models, while premium uses frontier models that cost several times more to run. At the start of benchmark creation and each benchmark run, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for the tokens actually used, regardless of the estimate. Grading works the other way round: the rates above are a price, not an estimate. We hold the quoted total for the whole pass, then charge the rate for each criterion or match actually judged — so a grading pass never costs more than its quote, and costs less when it judges fewer units than planned. Find the best model for your work Bring your own tasks or describe your use case, and get graded results with the costs attached. Try Optima