2026-10-11 16:37 UTC

Artificial Analysis claims its available Optima service builds and grades custom benchmarks from users' tasks and data across models and external agents, enabling workload-specific selection using measured quality, cost, and execution time.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumevaluation llm-tooling inference-economicsArtificial Analysis

What is this?

Artificial Analysis is described in the supplied third-party snippets as an independent benchmarking platform comparing AI models and API providers on quality, speed, latency, and price. The case attributes Optima to Artificial Analysis and claims it builds and grades benchmarks from users’ tasks and data, supports cross-model runs and external agents over HTTP, and reports per-task cost and execution time. None of the supplied web snippets mentions Optima, so its availability and specific capabilities remain unverified here; the custom-evaluation features described in the LayerLens result belong to a different company.

Why it matters to Scott

Artificial Analysis’s claimed Optima offering converges with Scott’s Capability Audit position and his trace-backed agent comparison work: representative-task evaluation across models and external agents could provide a build-versus-buy option for his backend benchmarks and evidence for task-aware routing. This is a concrete tooling connection rather than merely another endorsement of evaluation, but Optima’s availability and capabilities remain unverified in the supplied snippets, and the claims do not establish Scott’s requirements for preserved traces, calibrated controls or disclosed harnesses.
ip:concept.capability-auditip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisondev:project.remote-execdev:concept.task-aware-model-routingradar:concept.agent-evaluationradar:concept.model-evaluationradar:concept.model-routingradar:agent-review-studio-local-evaluation
queries asked of Scott's wikis
  • workload-specific evals versus public benchmark model selection
  • agent harness evaluation custom endpoints HTTP
  • task rubrics automated grading reliability
  • model routing quality cost latency tradeoffs
  • internal evaluation tooling build versus buy

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 727h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-11 09:28 (minted)⭐ origin echo-reconstructedOptima offers task and rubric construction, cross-model runs, external-agent participation over HTTP, and grading with per-task cost and tim
Artificial Analysis on blog (echo) · attributed from hn.story.49655513 · published time unknown
—
09-11 09:14first on hacker news · published · lag ?Build your own custom benchmark
prathje
—
09-11 09:14amplified on hacker newshn.story.49655513
prathje
peak 1 · 1 comments · 33% of case engagement
09-16 14:25amplified on hacker news 👑hn.story.49727542
fortitudedev
peak 3 · 1 comments · 65% of case engagement
09-11 09:21our radar first saw it · lag ?discovery anchor: hn.story.49655513—
pace: p43 vs 519 stories at the 720h mark (now 727h old) — ahead of checkly-agentic-go-rewrite (1.2x), behind api-delta-manifest (0.8x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBuild your own custom benchmark
Retrieved article excerpt

Open article · Retrieved 2026-09-11T09:23:06.449895+00:00

Artificial Analysis K Optima Build your own custom benchmark Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency Try Optima How it works Contract Review Benchmark Build agent drafting tasks Benchmark how well models review our supplier contracts — flag risks, extract terms. Working Reading your example contracts Drafting grading rubrics Writing task 18 of 24 Contract Review Benchmark 24 tasks · rubrics attached · 4 categories 1 . Build 2 . Run 3 . Grade Why build your benchmark with Optima Results specific to your use case Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released. Cut cost and time by over 10x Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models Built on Artificial Analysis grading expertise Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head How it works 1 Give Optima your context Describe it Import your traces Use your coding agent Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you. Describe your use case | Attach examples pptx AA-Merch_Sales_Pitch_Deck pptx AA-Merch_Startup_Swag_Pitch two decks we were happy with — this is what good looks like Build agent drafts tasks + rubrics Your drafted benchmark Sales Deck Creation 5 tasks · 10 criteria · head-to-head Northstar Hotels welcome-kit pilot Crescent Arts Museum collection pitch Apex Trails 12-month merch programme Tidepool summer pre-season launch City Sound Festival activation 2 Choose your evaluation type Task style Objective Rubric judge Subjective Optima Q&A The model answers questions that have known correct answers Document input The model answers questions about your uploaded files Agentic The model completes tasks and produces deliverables, using tools along the way Interaction Simulate a real conversation with different types of users Coming soon Q&A Objective grading Example task prompt Which HS tariff code applies to lithium-ion e-bike batteries? How it's graded The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion. 3 Run across models Your benchmark 24 tasks Flag the risky clauses Extract the payment terms Draft the counterparty summary same tasks, same conditions, every model Run sandbox + tools Models you picked Claude Fable 5 GPT-5.6 Sol Kimi K3 Gemini 3.6 Flash tokens, cost and time recorded per task 24/24 Or bring your own agent Your own agent can compete in the same run, over HTTP. Claude Opus 4.6 GPT-5.2 Gemini 3 Pro your-agent POST /runs/9f2c/artifacts Benchmark run 24 tasks · same judges Leaderboard 1 Claude Opus 4.6 3 GPT-5.2 4 Gemini 3 Pro 2 your-agent 4 Grade and decide Contract Review Benchmark 24 tasks · 4 models Model Score Cost / Task Time / Task Claude Fable 5 60 $ 0.18 74 s GPT-5.6 Sol 59 $ 0.11 52 s Kimi K3 57 $ 0.04 61 s Gemini 3.6 Flash 50 $ 0.02 29 s strong and cheap Score Cost per task $ 0.00 $ 0.05 $ 0.10 $ 0.15 $ 0.20 Claude Fable 5 60 · 74 s per task GPT-5.6 Sol 59 · 52 s per task Kimi K3 57 · 61 s per task Gemini 3.6 Flash 50 · 29 s per task Pricing Building and running a benchmark is priced on token usage. Grading is priced per unit judged. Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top. Rubric grading is $0.002 per criterion, per model with standard judges, or $0.040 with premium judges. Pairwise grading is $0.006 per match with standard judges, or $0.150 with premium judges. Rubric grading uses one judge from the selected tier; pairwise grading uses the full judge panel. Standard uses high-capability models, while premium uses frontier models that cost several times more to run. At the start of benchmark creation and each benchmark run, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for the tokens actually used, regardless of the estimate. Grading works the other way round: the rates above are a price, not an estimate. We hold the quoted total for the whole pass, then charge the rate for each criterion or match actually judged — so a grading pass never costs more than its quote, and costs less when it judges fewer units than planned. Find the best model for your work Bring your own tasks or describe your use case, and get graded results with the costs attached. Try Optima
prathje11
🟧 echo.blog ⭐Optima offers task and rubric construction, cross-model runs, external-agent participation over HTTP, and grading with per-task cost and timArtificial Analysis——
🟧 hnShow HN: OmnisBench, a re-gradable, open LLM routing benchmark on fresh tasksfortitudedev30

Interpretation history

Decision trace