2026-10-11 17:11 UTC

Aven-1 creator neurovance claims the released repository implements a 20M-parameter LLM from scratch including PPO-based RLHF, potentially providing builders a small-scale reference for the full training pipeline.

state: expiredheat: lowuncertainty: highnovelscott: lowsmall-model-training rlhf open-llm-artifactsneurovance

What is this?

The supplied case identifies Aven-1 as a repository released by a creator using the handle neurovance, whose Hacker News submission claims a 20-million-parameter LLM built from scratch, including PPO-based RLHF. None of the supplied search-result snippets directly documents Aven-1, its creator, or its repository, so the implementation, training completion, and usefulness as a full-pipeline reference remain unverified. The web answer's additional claims about preference-based exploration and application design are not supported by Aven-1-specific evidence in the supplied results.

Why it matters to Scott

Scott’s Snake DQN lab establishes an interest in inspectable, hands-on reinforcement-learning experiments, but the hits establish no PPO/RLHF project or claim that Aven-1 would affect; this is a possible learning reference, not a demonstrated reason to change what he builds or argues. The radar tracks related training references but not this development, and Aven-1’s claimed full pipeline remains unverified.
dev:project.snakeradar:trip-plain-c-transformer-stackradar:nanorl-lightweight-llm-rl-trainerradar:concept.small-modelsradar:concept.model-training
queries asked of Scott's wikis
  • small-model training from scratch educational reference implementations
  • PPO RLHF reward modeling preference data training pipelines
  • reproducible open-model artifacts training code checkpoints evaluations
  • local model training compute budgets versus fine-tuning
  • hands-on model internals learning projects

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnI built a 20M-parameter LLM from scratch, including full RLHF (PPO)neurovance10
🟧 echo.github ⭐The creator's HN submission presents Aven-1 as a 20M-parameter LLM built from scratch, including full RLHF using PPO.ariveinberg-max——

Interpretation history

Decision trace