2026-10-11 16:37 UTC

Deforget developer Kaloyan Lachezarov reports silent Apple Foundation model changes throughout the iOS 27 rollout, including within unchanged OS builds, making OS-version pinning insufficient for reproducible on-device inference.

state: watchingheat: mediumuncertainty: mediumknownscott: lowlocal-inference apple-on-device-models evaluation-drift agent-harnessesKaloyan LachezarovDeforgetApple

What is this?

Apple provides developers access to its on-device language model through the Foundation Models framework, according to its supplied research snippet; other results describe iOS 27 and its Apple Intelligence updates. The case attributes to Deforget developer Kaloyan Lachezarov a longitudinal evaluation reporting model behavior changes across iOS 27 betas, including silent downloads under unchanged OS build numbers. None of the supplied search results identifies Lachezarov or contains that evaluation, so the central claim—and its implication that OS-version pinning cannot ensure reproducible inference—remains unverified here.

Why it matters to Scott

The radar already tracks Apple's silent on-device model changes in radar:apple-silent-on-device-model-changes; changes within unchanged OS builds would sharpen that story, but this additional claim remains unverified in the supplied material. It connects to Scott's Drift Monitoring and Evaluation-Driven Development positions, but no hit establishes a Foundation Models dependency or a verified consequence requiring him to change his harnesses.
ip:concept.drift-monitoringip:concept.evaluation-driven-developmentradar:apple-silent-on-device-model-changes
queries asked of Scott's wikis
  • local inference model ownership and update control
  • model version pinning reproducible inference
  • evaluation drift regression testing agent harnesses
  • Apple Foundation Models on-device app projects
  • model artifact provenance behavioral change monitoring

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 1826h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

07-27 14:00⭐ origin echo-reconstructedThe living evaluation report describes behavioral changes across iOS 27 betas and silent downloads under unchanged build numbers; its Septem
Kaloyan Lachezarov on blog (echo) · attributed from hn.story.49710030
—
09-15 09:50first on hacker news · published · +1195.8hI measured Apple's on-device LLM across an OS beta cycle
lachezarov
—
09-15 09:50amplified on hacker news 👑hn.story.49710030
lachezarov
peak 2 · 0 comments · 98% of case engagement
09-15 10:20our radar first saw it · +1196.3hdiscovery anchor: hn.story.49710030—

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnI measured Apple's on-device LLM across an OS beta cycle
Retrieved article excerpt

Open article · Retrieved 2026-09-15T10:22:07.394786+00:00

Engineering notes

# I measured Apple's on-device model so you don't have to

*July 28, 2026 · Kaloyan Lachezarov*

*This is a living post. The corpus reran on every iOS 27 beta, on the release candidate, and on the public build. Last updated September 14, 2026, launch night: the public build (24A437) is newer than the release candidate and carries a different model - quieter on people and on dated sentences - and the corpus caught it within hours: raw 43 percent, final 82, restraint 34 of 35. The gate still passes. The three newest chapters cover beta 8, the release candidate and the public build, and how the dice got taken out of the instrument.*

Deforget turns diary prose into reminders, calendar events, meeting notes, and an index of people - using Apple's on-device Foundation model, with no cloud fallback, because [nothing leaves the device](https://www.deforget.app/blog/on-device-diary). That choice has a consequence people underestimate: **you cannot pin the model.** It belongs to the user's operating system. Whatever iOS ships, my users get, and every OS update can move the ground under the product.

When you depend on a model you can't version, you need instruments, not vibes. So Deforget carries its own evaluation harness: a fixed corpus of diary scenarios with typed expectations - dates, events, reminders, commitments, meeting notes, people, decisions, deduplication, and restraint (emotional prose must extract *nothing*). It runs on real iPad hardware against every iOS beta, and every case is scored twice: once on the **raw** model output, and once on the **final** result after a deterministic repair layer. The gap between those two numbers is the engineering, measured.

## The numbers

Method notes that matter: OS comparisons use the same app binary - build once, run on the old OS, update the OS, run again without rebuilding, so nothing moves but the model. Five runs per case, because single runs lie. The corpus grows as field reports arrive, so totals shift between revisions; compare percentages, and only trust a raw-to-raw comparison within one corpus version.

| OS build | raw model | after repair layer |
| --- | --- | --- |
| **iPadOS 26.5** | 46% (78/170) | 75% (127/170) |
| Mood prose became junk commitments (restraint passed 3 checks of 25); 9 of 170 model calls died in runaway-generation loops | | |
| **27.0 beta 3** | 61% (104/170) | 91% (155/170) |
| Same binary as the row above - only the model changed. Restraint 21/25, zero call failures, meetings and dedupe perfect | | |
| **27.0 beta 4** | 59% (106/180) | 89% (161/180) |
| Larger corpus (two new adversarial cases, one deliberately raw-hostile - the raw dip is composition, not regression). Restraint 25/25, dates 25/25, still zero failures | | |
| **27.0 beta 5** | 47% (92/195) | 63% (122/195) |
| The regression this method exists for: restraint collapsed to 22 of 35 checks, and dates, events, and decisions collapsed toward emitting nothing. Nothing in my pipeline had changed | | |
| **27.0 beta 5, one night later** | 48% (93/195) | 82% (160/195) |
| Same OS, same model - the raw column agreeing is the proof. Every recovered point is deterministic: restraint guards first, then a floor under recall. Restraint 33/35 | | |
| **27.0 beta 6** | 46% (91/200) | 83% (166/200) |
| Corpus grew again (200 checks). A fresh beta 5 anchor on the identical binary and corpus ran hours before the update: raw 47%. The two raw columns agreeing means the model is frozen at beta 5 - and the launch gate passes on it. Restraint 33/35, zero failures | | |
| **27.0 beta 6, new model** | 43% (88/205) | 83% (171/205) |
| Same OS build (24A5418b), re-measured on August 22. Beta 6's model had not arrived with the update - it turned up silently later: decisions collapsed back toward emitting nothing, feelings crept into the output again, and two same-evening runs disagreed by three raw points - the new model is noisier too. The repair layer held final at 83% in both | | |
| **27.0 beta 7** | 42% (86/205) | 85% (174/205) |
| Arrived August 24 carrying beta 6's late model unchanged - raw inside the band, every failure a known shape. Restraint back to 33/35 with margin; the launch gate passes | | |
| **27.0 beta 7, new model** | 48% (99/205) | 86% (176/205) |
| Same OS build (24A5424a): another model arrived by download between August 25 and 27 - and this one is better. The first raw improvement since beta 5, the feelings junk gone from the raw output, and the best final score of the series. Three runs land between 48 and 50% raw - a tighter model, too | | |
| **27.0 beta 8** | 47% (97/205) | 85% (174/205) |
| Arrived August 31 carrying late beta 7's model - raw inside the band, restraint final 34/35, the launch gate passing with room. The build string (24A5430a) then held for nine days while the model under it moved three more times; the row below is the low point | | |
| **27.0 beta 8, September 3** | 40% (82/205) | 81% (167/205) |
| A third silent swap in three days: quiet on dated sentences and junky on decisions again. Restraint 29/35 - the only run since beta 5 under the launch gate's restraint bar; a confirmation run the same night read 33, so this model sat on the bar rather than beneath it. The floor under dated recall did the carrying | | |
| **27.0 (24A435), release candidate** | 49% (100/205) | 86% (176/205) |
| September 9, five days before the public release. Beta 8's original model is back, dead centre of its band; restraint 35/35, zero call failures, every miss a known shape. This is the launch run | | |
| **27.0 (24A437), the public build** | 43% (89/205) | 82% (168/205) |
| Launch night, September 14. Apple shipped a build newer than the candidate, and the corpus says it carries a different model: raw back down to the quiet early-September shape, people mentions down by a third, dated recall down, restraint clean at 34/35. Two runs agree (raw 90 and 89). The gate passes; the margin over iOS 26 is two points | | |

The number I watch most isn't in any single row: the repair layer's contribution. It stayed close to thirty points while the model improved - 29 on iOS 26, 30 on beta 3, 31 on beta 4 - then jumped to 34 the night the model regressed, 37 on beta 6, 40 against the model that later slipped in under the same build number, and 43 on beta 7's carried-over model. Then the better model landed and the contribution fell back to 38 - the first time the number has shrunk, and it shrank because raw rose to meet it. Beta 8 held it at 38, the September 3 low point pushed it to 41, the release candidate settled at 37, and the public build put it back at 39. That asymmetry is the design. When Apple ships a better model, the layer doesn't shrink; it stands on higher ground. When Apple ships a worse one, the layer carries more of the ceiling instead of falling with it.

## What Apple actually fixed in iOS 27

- **Restraint.** The worst failure mode on iOS 26 was "I feel really good about how today went" becoming a *commitment*. Feelings turning into to-dos is the fastest way to make someone turn intelligence off - or stop writing honestly, which is worse. Across the first three 27 betas the restraint category went from 12 percent of checks to 84 to **100**. (Beta 5 then broke it again - that story is below, and the repair layer now holds the line the model dropped.)
- **Reliability.** On 26, roughly one call in twenty entered a runaway loop - the model babbling until it blew its own context window. On every 27 build since, the release candidate included: zero, across thousands of calls.
- **Vocabulary.** The *decision* type effectively didn't exist in 26's output; everything flattened into commitments. 27 emits it unprompted.
- **Consistency.** 26's failures were flaky - one run in two. 27's failures are consistent, and a consistent failure is something you can engineer around. A flaky one isn't.

## What's still broken, and what I do about it

Three failure classes survived beta 4, and each gets a different treatment - that's the point of measuring instead of guessing.

- **Type confusion at the commitment/event boundary.** "The offsite is on the 14th" comes back as a commitment, not a calendar event - consistently. After the third consecutive beta reproduced it, a pre-agreed deterministic retype rule went in: dated, scheduled phrasing outranks the model's label. That rule was written weeks earlier and sat disabled until the reports met its conditions. Fixes here follow a playbook, not a mood.
- **Detached clock times.** "Next Thursday's sync at 3:30" once resolved the time separately from the day and produced 3:30 *AM*. The repair layer now grafts a detached clock onto the sentence's date before it can misfire. Deterministic, tested, no model involved.
- **Recall flakes.** One decision case sometimes returns nothing at all. The discipline for missing output is different: wait. Across this beta cycle, recall has recovered on its own while precision held - and a silent miss is recoverable in the app, while junk output is trust-fatal. Precision gets engineering; recall gets patience. (Patience, it turned out, has a floor - beta 5 found it.)

## Beta 5 took the ground away

Everything above described a model improving beta over beta. 27.0 beta 5 (24A5408d) reversed it. The first corpus run after the update posted the worst restraint since iOS 26 - hedged musings came back as *decisions*, past narrative ("We walked along the river...") came back as commitments, mood prose got stamped with phantom event dates - while dates, calendar events, and decisions collapsed toward emitting nothing at all. Final fell to 63 percent; beta 4 had stood at 89. Nothing in my pipeline had changed between those runs. The raw column is the receipt.

This is the scenario the regression playbook was written for, so the night ran on rails instead of adrenaline, in the pre-agreed order. Restraint first, because junk is trust-fatal and silence is merely disappointing: each new junk class got a deterministic guard - past-tense openers stop becoming plans, the hedge rule now covers decisions, an event dated by nothing but the prompt's own anchor dies. Then recall. Waiting is the usual answer for a model that says nothing, but a collapse this deep has a pre-agreed answer too: a deterministic floor that acts only when the model returned nothing actionable and the sentence itself names a future date plus an unambiguous frame - "remind me", a meeting word backed by a clock time, a first-person modal, "by Friday". No date, no entity, so the floor cannot regress restraint by construction. A past date is a memory, not a plan, and stays one.

I reran the corpus after every change and stopped when the gate passed. The two beta 5 rows in the table are the first and last runs of that night. Between them the raw score barely moved - same OS, same model, no prompt changes - which is exactly the point: every recovered point is inspectable, testable code, not a plea to the model. What the floor deliberately does not do is also on the record: sentences with no date stay silent even when recall is hurting, because a diary that invents structure is worse than a diary that misses some.

## Beta 6 held still

27.0 beta 6 (24A5418b) arrived on August 17 with a tell: the update pulled down no new AI assets, and the app's own model-status check reported the Foundation model ready the moment the OS came back up. That usually means the model didn't change. But a suspicion isn't a measurement, so the corpus ran twice that night - first a fresh beta 5 anchor on the exact binary and corpus the comparison would use, then beta 6, with nothing moved but the OS.

Raw went 47 to 46 percent, noise-level agreement, and every failure landed in an already-known shape. The model is frozen at beta 5. The repaired pipeline posted 83 percent with restraint at 33 of 35 and zero call failures - both launch-gate conditions pass. A couple of decision cases even came back on their own, which is
lachezarov20
🟧 echo.blog ⭐The living evaluation report describes behavioral changes across iOS 27 betas and silent downloads under unchanged build numbers; its SeptemKaloyan Lachezarov——

Interpretation history

Decision trace