The case describes “Bart,” a 2.82B-parameter model allegedly trained from scratch by Unbounded Labs on 20.1B tokens of pre-1931 English, intended for experiments in historical-language behavior, idea generation, and controlled-data evaluation. However, the supplied search results instead document “Talkie,” a different 13B-parameter open-weight model associated with Alec Radford, Nick Levine, and David Duvenaud and likewise trained on pre-1931 text. The snippets therefore support the broader research pattern—historically bounded corpora as low-contamination testbeds—but do not independently establish Bart’s release, specifications, ownership, or claimed research value.
Scott’s Future-Leakage Rule already captures the core position: temporally freeze available knowledge to avoid contamination during historical evaluation. Bart would be another possible implementation of that principle, but the supplied material neither establishes the model’s release nor provides evaluation results that would extend or challenge Scott’s position.
ip:concept.future-leakage-ruleip:concept.cognitive-provenanceradar:concept.model-evaluationradar:concept.benchmark-integrityradar:concept.open-model-trainingradar:concept.small-models
queries asked of Scott's wikis
- historically bounded training data and contamination-free evaluation
- small language models as controlled experimental testbeds
- training-data cutoffs and knowledge provenance
- novel idea generation versus memorization
- open-weight models for scaling-law research
- dataset curation and temporal knowledge boundaries
2026-08-27T04:24:05Z
After more than 48 hours, no controlled evaluation, provenance test, or substantive third-party use has emerged; repeated sensor activity was only engagement churn. Bart remains available for future rediscovery, but this release has not developed within its active monitoring horizon.
2026-08-25T03:31:34Z
The refreshed discussion adds no reproducible provenance test, independent evaluation, or third-party implementation; it remains repetitive criticism of naming and anecdotal output quality. Bart is still a publicly testable but unvalidated example of temporal data bounding.
2026-08-25T01:23:51Z
The refreshed comments repeat existing naming, prior-art, and poor-output critiques without adding a reproducible contamination test or independent evaluation. Bart remains a testable but unvalidated implementation of an already-known temporal-bounding pattern.
2026-08-24T23:30:05Z
The refreshed discussion remains repetitive amplification of naming confusion, Talkie prior art, and anecdotal poor outputs; it adds no controlled provenance test or independent evaluation. Bart remains publicly testable but has not earned stronger research significance or corroboration.
2026-08-24T20:47:03Z
Refreshed discussion adds more anecdotal evidence of weak factual quality and highlights prior art in Talkie, but still provides no controlled contamination test or independent evaluation. The case remains a testable but unvalidated implementation whose research value is uncertain.
2026-08-24T18:27:36Z
Public weights and a live demo make Bart a testable release rather than merely an announcement, warranting passive monitoring. The second post is still first-party amplification, while early outputs raise unresolved provenance and basic-quality concerns without establishing either contamination or research utility.
2026-08-24T18:23:05Z
evidence attached: reddit.post.1vx94er — The post corroborates that Bart has been released with weights, a demo, and training details for independent evaluation.
2026-08-24T17:27:29Z
A user’s report that Bart referenced Mr. T introduces the first concrete challenge to the claimed pre-1931 training boundary, but a single prompt result cannot distinguish contamination from hallucination or indirect period-compatible associations. The case now has a provenance question rather than independent evidence of research utility.
2026-08-24T16:29:33Z
The modest engagement increase is amplification of the original announcement, not independent validation. Without evaluations, third-party use, or confirmation of the model and corpus claims, Bart remains an unproven implementation of an already-known temporal-bounding pattern.
2026-08-24T16:26:59Z
grounded: known/low — Scott’s Future-Leakage Rule already captures the core position: temporally freeze available knowledge to avoid contamination during historical evaluation. Bart
2026-08-24T16:24:25Z
case created — The announcement exposes a usable model and demo with an unusual, independently testable training-data constraint, but currently has only one low-engagement observation.