2026-10-11 17:11 UTC

Indic ModernBERT creator kkkamur claims to have trained a released 188M-parameter Hindi-first encoder with 8,192-token context on roughly 28.5 billion tokens using one RTX 4090 in about five days, potentially making long-document Hindi retrieval models practical to develop on consumer hardware.

state: expiredheat: lowuncertainty: highnovelscott: lowmultilingual-retrieval rag open-modelskkkamur07

What is this?

The Hugging Face release kkkamur07/hindi-modernbert is a Hindi adaptation of ModernBERT, with a new tokenizer and approximately 28 billion tokens of Hindi pretraining according to its release snippet. The case attributes claims of 188 million parameters, an 8,192-token context window, and training on one RTX 4090 in about five days to its creator; the supplied release snippet does not independently establish those specifications or the hardware/time claim. ModernBERT itself is a separate encoder family announced by LightOn with Answer.AI and collaborators, whose supplied descriptions support 8,192-token inputs and retrieval applications, but do not establish this Hindi release’s retrieval quality or training economics.

Why it matters to Scott

Scott uses BGE-M3 for multilingual advisory recall, but the hits establish neither a Hindi retrieval requirement nor a language-specific encoder-training programme; this release does not yet warrant changing his embedding stack because its retrieval quality and consumer-GPU training economics remain unverified. The radar tracks related consumer-GPU training claims, but no supplied page tracks this Hindi release, and it does not substantively converge with or challenge a Scott position.
dev:technology.bge-m3radar:concept.multilingual-modelsradar:concept.model-trainingradar:bananamind-2-pro-consumer-gpu-training
queries asked of Scott's wikis
  • consumer GPU training economics language-specific encoders
  • Hindi multilingual retrieval evaluation RAG
  • long-document embeddings chunking retrieval quality
  • open model adaptation custom tokenizers domain pretraining
  • local knowledge systems embedding model selection

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: 188M Hindi encoder, 28B tokens, 8K context, 1Γ— RTX 4090kkkamur10
🟧 echo.github ⭐The creator links this repository for a Hindi-first ModernBERT trained from scratch, reporting 188M parameters, approximately 28.5B Hindi tokkkamur07β€”β€”

Interpretation history

Decision trace