TDQS’s maintainer claims the released scoring specification can quantify MCP tool-definition quality and guide schema improvements, potentially standardizing how agent tool discoverability and selection are assessed.
state: expiredheat: lowuncertainty: highknownscott: mediummcp tool-discovery agent-harnessespunkpeyeTDQSGlama
What is this?
TDQS (Tool Definition Quality Score) is an open-source framework published by Glama and attributed to maintainer punkpeye for evaluating MCP tool definitions across six dimensions: purpose clarity, usage guidelines, behavioral transparency, semantics, conciseness, and completeness. This targets MCP’s model-controlled, dynamically discoverable tools, where definition quality may affect an agent’s ability to select and invoke the right tool. Glama claims the score can guide schema improvements and provide a standardized assessment, but the supplied snippets do not establish independent validation, adoption, or a demonstrated link between TDQS scores and agent task success.
Why it matters to Scott
The radar already tracks this problem on `radar:mcp-server-agent-usability`, with TDQS supplying a specific static scoring proposal rather than establishing a new result. It bears directly on Scott’s MCP connector and could serve as the inspection rung of his Progressive Evaluation Ladder, but its value remains unproven until scores are correlated with retained agent paths, tool-selection accuracy, and task success as required by Reflexive Agent Design.
ip:framework.reflexive-agent-designip:concept.progressive-evaluation-ladderdev:project.mcp-ip-wikiradar:mcp-server-agent-usabilityradar:oqoqo-agent-interface-evalsradar:concept.agent-evaluation
queries asked of Scott's wikis
- tool schema quality and agent tool selection
- MCP tool discoverability evaluation
- agent harness tool-definition linting
- tool descriptions versus execution success
- automated schema improvement loops
- standards for agent-facing tool interfaces
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-09-07T19:37:50Z
The launch episode has faded without new evidence of adoption or predictive validity, and no concrete follow-up is expected. TDQS remains a possible inspection rubric for Scott’s MCP tooling; expiration ends active tracking of this episode, not the possibility that the rubric proves useful.
2026-09-05T18:31:53Z
TDQS remains a concrete inspection rubric for Scott’s MCP tooling, not evidence of a predictive quality metric or emerging standard. This look adds no substantive evidence of adoption or task-outcome validation; the reported comment is unavailable and cannot establish either.
2026-09-03T18:00:24Z
No independent validation, implementation uptake, or evidence linking TDQS scores to agent task outcomes has appeared; the small engagement increase does not change the case’s meaning. It remains a plausible static inspection rubric rather than an established MCP quality standard.
2026-09-03T17:46:37Z
grounded: known/medium — The radar already tracks this problem on `radar:mcp-server-agent-usability`, with TDQS supplying a specific static scoring proposal rather than establishing a n
2026-09-03T17:43:01Z
case created — The open specification targets a recurring MCP interoperability failure with a concrete scoring mechanism from an established ecosystem contributor.
Decision trace
- 09-08 05:37expireThe launch episode has faded without new evidence of adoption or predictive validity, and no concrete follow-up is expected. TDQS remains a possible inspection rubric for Scott’s MCP tooling; expirati
- 09-08 05:37alert_silentThis look contains no new consequential delta or pending confirmation. The previously established specification release can remain briefing material without another interruption.
- 09-08 05:37alert_routeThis look contains no new consequential delta or pending confirmation. The previously established specification release can remain briefing material without another interruption.
- 09-06 04:31repriceTDQS remains a concrete inspection rubric for Scott’s MCP tooling, not evidence of a predictive quality metric or emerging standard. This look adds no substantive evidence of adoption or task-outcome
- 09-06 04:31alert_silentThe maintainer’s specification release is established, but this delta contains only engagement changes. There is no new tooling capability, adoption evidence, or actionable result that makes the next
- 09-06 04:31alert_routeThe maintainer’s specification release is established, but this delta contains only engagement changes. There is no new tooling capability, adoption evidence, or actionable result that makes the next
- 09-04 04:00repriceNo independent validation, implementation uptake, or evidence linking TDQS scores to agent task outcomes has appeared; the small engagement increase does not change the case’s meaning. It remains a pl
- 09-04 04:00alert_silentThe only new signal is minor engagement without discussion or substantive evidence, so there is no consequential delta that warrants attention before the next briefing.
- 09-04 04:00alert_routeThe only new signal is minor engagement without discussion or substantive evidence, so there is no consequential delta that warrants attention before the next briefing.
- 09-04 03:54alert_silentA known MCP ecosystem maintainer has released a concrete open-source rubric for scoring tool definitions, but the supplied evidence shows neither adoption nor validation that TDQS scores predict tool
- 09-04 03:54surface_candidateA known MCP ecosystem maintainer has released a concrete open-source rubric for scoring tool definitions, but the supplied evidence shows neither adoption nor validation that TDQS scores predict tool
- 09-04 03:54alert_routeA known MCP ecosystem maintainer has released a concrete open-source rubric for scoring tool definitions, but the supplied evidence shows neither adoption nor validation that TDQS scores predict tool
- 09-04 03:46groundThe radar already tracks this problem on `radar:mcp-server-agent-usability`, with TDQS supplying a specific static scoring proposal rather than establishing a new result. It bears directly on Scott’s
- 09-04 03:43createThe open specification targets a recurring MCP interoperability failure with a concrete scoring mechanism from an established ecosystem contributor.