2026-10-11 16:36 UTC

agentic-coding

band: warmmomentum: stable score: 0.454
temperature history

Episodes (10)

Independent evaluations will determine whether LG AI Research's Apache-2.0-licensed K-EXAONE 2.0 750B-A37B delivers competitive multilingual, coding, long-context, and tool-use performance among leading open-weight models.
expired
Linear claims its CI redesign roughly halved runner time per test and reduced PR waits from over six minutes to just over five despite an almost fourfold increase in test suites, demonstrating a practical response to agent-driven validation load.
resolvedconvergesscott: medium
IQuest claims its released IQuest-Q1 β€” a 320B-total/~15B-active MoE purpose-built for agentic coding, reasoning, and multi-step tool use β€” is a capable open-weight coding-agent model; community adoption and independent measurement of it in local coding-agent workflows will determine whether it earns a practical place or fades as another unreleased-in-practice announcement.
watchingconvergesscott: medium
Microsoft claims its released FrogNano-4B-2609 β€” Qwen3.5-4B post-trained with RL on synthetic repository-level SWE environments for the Leaf five-tool harness β€” makes practical agentic coding viable on GPU-poor local hardware; sustained community adoption (bartowski GGUFs already exist), independent benchmark results in real agent harnesses, or quiet fading resolves whether a 4B open model becomes a credible local coding-agent default.
seedconvergesscott: medium
Reddit user ai_art_is_art claims agent-assisted, 100% Rust clean-room replacements of seven Adobe apps (Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign, Acrobat Pro) are released as functional open-source software with apps and code published; usable repositories and working apps confirm agent-built full-application suites as a real capability marker, while vaporware, broken releases, or debunking closes it as overclaim.
resolvedconvergesscott: low
DHH claims agents implemented and optimized the entire Campfire web app in Elixir, Go, and Ruby-replacing ports β€” conceding Ruby is slowest and asking 'if you're no longer reading the code?' β€” and whether the three ports in basecamp/once-campfire hold up under inspection, or other prominent builders replicate whole-app agentic porting, establishes whole-application agentic implementation as working practice rather than provocation.
watchingconvergesscott: high
Redditor MushroomMan234 reports that UkisAI's Swift 1.5 β€” a reasoning-efficient fine-tune of Qwen3.8-Flash-Next β€” running on sf-stav's veloGB10 GB10-only engine sustains ~110 tok/s decode on two DGX Sparks (vs ~52 for base NVFP4 on vLLM) and beats base Flash-Next at medium effort on coding-agent pass rate (92% vs 50%), making a fine-tune-plus-single-model-engine stack a demonstrated local coding-agent path on GB10 hardware if others replicate it.
seedconvergesscott: medium
Mupt AI claims SelfBench β€” which converts a repository's merged PRs into Harbor-gated tasks with hidden tests and publishes accuracy-vs-cost leaderboards β€” becomes a standard private gate teams use to benchmark coding agents on their own codebases; external teams running it and releasing results confirm it, quiet fade closes it.
seedconvergesscott: medium
Two same-day independent builders claim Claude-driven agent pipelines produced professional animated video β€” one fully code-rendered via Remotion with 58 sub-agents in 38 hours, one via an MCP-connected video editor in 2-3 prompts β€” and whether further builders replicate the pattern or the demos stay one-offs resolves it.
resolvedconvergesscott: high
Firelex claims his released Jeff-Code β€” a 0.8B decision model inside Pi's agent loop that takes routine steps itself and routes Qwen 3.8-27B's thinking β€” cuts time per coding task to 0.68x at an unchanged pass rate across 1,242 paired benchmark tasks, and independent replication or adoption makes small-decision-model-in-the-loop a standard coding-agent acceleration.
seedconvergesscott: high

Trajectory notes