2026-10-11 16:37 UTC

GitHub claims its new unified Copilot inline model replaces separate completion, nearby-edit, and long-distance-edit models with multi-edit patch generation and caching, improving suggestion selection and reducing follow-up editing latency.

state: seedheat: mediumuncertainty: mediumconvergesscott: lowcoding-agents copilot llm-servingGitHubJulia GongBen LiggettUlugbek Abdullaev

What is this?

GitHub Copilot is GitHub’s coding assistant; its inline suggestions complete code and predict edits, including changes away from the cursor, according to GitHub’s documentation. The case describes a new unified model generating multiple diff patches and caching subsequent edits, but the supplied search snippets do not establish that architecture, its rollout, or the claimed selection and latency improvements. One GitHub documentation snippet still distinguishes completion-model selection from the model used for next-edit suggestions, leaving the relationship to the claimed unification unresolved; the snippets also do not establish the named individuals’ roles.

Why it matters to Scott

GitHub’s claimed generation and caching of subsequent edits converges with Scott’s Cognitive Prefetching principle—doing likely future work before it becomes latency-critical—and his GitHub account history records early Copilot adoption. However, the supplied grounding does not establish the architecture or gains, and the hits show neither current dependence on Copilot inline suggestions nor evidence challenging his context-sensitive editing claims; this remains a potential example of his pattern rather than an actionable development.
ip:concept.cognitive-prefetchingwork:project.githubradar:concept.inference-latencyradar:zed-paid-edit-predictions
queries asked of Scott's wikis
  • coding assistant inline completion versus next-edit prediction
  • diff patch generation multi-edit coding harnesses
  • LLM inference latency speculative generation output caching
  • unified models versus task-specific model orchestration
  • coding suggestion acceptance evaluation developer editing workflow

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 626h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-15 14:00⭐ origin echo-reconstructedGitHub describes a single 3-in-1 inline-suggestions model using a diff-patch output format, with subsequent patches cached to provide faster
Julia Gong, Ben Liggett, and Ulugbek Abdullaev on blog (echo) · attributed from hn.story.49738337
—
09-17 09:30first on hacker news · published · +43.5hBuilding the New GitHub Copilot Inline Suggestions Model: Part One
soheilpro
—
09-17 09:30amplified on hacker news 👑hn.story.49738337
soheilpro
peak 1 · 1 comments · 98% of case engagement
09-17 10:21our radar first saw it · +44.4hdiscovery anchor: hn.story.49738337—
pace: p23 vs 1032 stories at the 336h mark (now 626h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBuilding the New GitHub Copilot Inline Suggestions Model: Part One
Retrieved article excerpt

Open article · Retrieved 2026-09-17T10:22:29.992292+00:00

# Building the new GitHub Copilot Inline Suggestions Model: Part One

September 16, 2026 by [Julia Gong](https://linkedin.com/in/juliagong), [Ben Liggett](https://www.linkedin.com/in/ben-liggett), and [Ulugbek Abdullaev](https://github.com/ulugbekna)

*Completion-style ghost text, next edit suggestions near the cursor, and edits farther away were previously powered by separate models. We built one model for all three, and learned that the best results come from training, evaluation, and editor design evolving together.*

At GitHub Copilot, our mission is to support all development workflows, from AI-assisted coding using inline suggestions to agent-first software engineering in the VS Code Agents window. Inline suggestions are used and loved by millions of developers, and we continue to push the quality bar on it and all GitHub Copilot experiences.

Writing code is rarely linear. The next useful change might be a few characters at the cursor, a nearby rewrite, or a related edit elsewhere in the codebase. Inline suggestions accommodate these different workflows through completion-style ghost text, nearby next edit suggestions, and long-distance edits.

Previously, specialized model paths powered each type of suggestion. The tried-and-true [completions model](https://github.blog/ai-and-ml/github-copilot/the-road-to-better-completions-building-a-faster-smarter-github-copilot-with-a-new-custom-model/) doesn’t need much of an introduction, and in earlier posts, we shared how we [trained a custom model for next edit suggestions](https://github.blog/ai-and-ml/github-copilot/evolving-github-copilots-next-edit-suggestions-through-custom-model-training/) and [extended those suggestions to edits farther away](https://code.visualstudio.com/blogs/2026/02/26/long-distance-nes). We have now unified all three models behind a single “3-in-1” model that can do it all.

Unifying the models **improves suggestion quality** by allowing a **single model to choose the best edit for the developer’s current work end-to-end** instead of using programmatic logic to choose among specialized models. For instance, let’s say the developer has typed `class Fa` in the penguin feeding program below. Using the previous standalone models (left), the most mature and battle-tested completions model is always triggered first. Since it can only append to the prefix `Fa`, it does what it knows best and suggests `FastingPenguin`. However, in this particular case, a better NES suggestion exists, as suggested by the unified model (right): semantically rewrite `Fa` to `Fish`.

| Before (standalone models) | After (unified model) |
| --- | --- |
| A code example where the standalone completion model chooses FastingPenguin instead of the better Fish rewrite. | A code example where the unified model selects the better Fish rewrite for the same context. |

Not only can we tackle all existing tasks in one model, but because this model can **output multiple edits in one response**, we can cache these additional edits to deliver a **faster experience** for subsequent edits if the preceding suggestions were desirable. This creates a **smoother, snappier tab-tab-tab experience** across the editing flow.

To enable this better editing experience, we first **reframed the modeling task**—especially the prompt and output format—to generalize across code editing tasks. We then **combined our learnings in data quality, model training and evaluation, and client design** from the standalone models, in addition to extensive additional experimentation (over 200 offline and 20 online experiments, whose course we will chart in the remainder of these posts), to carefully optimize the model and end-to-end experience. The result is a **higher quality, faster, and more cohesive typing companion** that is greater than the sum of its parts.

An animation of a VS Code coding experience with ghost text and fast follow-up tab-tab-tab suggestions showing the unified inline suggestions experience.

## Why unify?

Prior to the unified model, the production implementation of inline suggestions in VS Code was powered by three separate models:

1. **Completions** (the [familiar ghost text](https://github.blog/ai-and-ml/github-copilot/the-road-to-better-completions-building-a-faster-smarter-github-copilot-with-a-new-custom-model/) we all know and love),
2. **NES** ([next edit suggestions](https://github.blog/ai-and-ml/github-copilot/evolving-github-copilots-next-edit-suggestions-through-custom-model-training/), or nearby edits bounded by a few lines above and below the cursor position),
3. **Long-distance NES** ([longer-range NES-style suggestions](https://code.visualstudio.com/blogs/2026/02/26/long-distance-nes) farther from the cursor).

This dedicated work on individual tasks, with a highly mature completions model alongside newer editing models, resulted in a client-orchestrated experience that triggered each where appropriate.

While effective for these individual tasks and a necessary step to bring new editing capabilities to users, this created key shortcomings not only in terms of user flow quality (sometimes making suboptimal or piecemeal edits rather than a single clear edit that best serves the context) and latency (potentially up to 4 model calls per opportunity, as shown in the figure below), but also in terms of extensibility to future features (for example, multi-file edits).

Standalone models diagram showing separate completion, NES, and long-distance NES call paths.

*Figure 1. The original stepping-stone production setup for VS Code inline suggestions with three standalone models: completions, NES, and long-distance NES. Note that this diagram is simplified for clarity and represents the behavior from a user’s perspective (real provider calling behavior is more complex and involves issuing parallel calls to completions and NES, for instance, depending on the position of the user’s cursor and how quickly the prioritized model responds). Each model handled its own standalone task, while the client orchestrated the logic of when to trigger which model. The completions model would always be triggered first to predict ghost text at the user’s cursor, which would be shown if ghost text was generated. If no ghost text was produced, the NES model would be invoked to rewrite the window around the user’s cursor, and shown using the appropriately rendered view kind if the rewrite resulted in a net edit. If still no edit was predicted, the long-distance edit model was triggered to predict a potential line number to jump to. If no line number was produced, no edit would be shown. If it did jump to a new line number, the NES model would be invoked once more to rewrite a window around the long-distance line number, resulting in a long-distance NES edit.*

Unification also creates a compounding benefit. With **one shared model powering every inline suggestion**, each model improvement or new model feature can now benefit the full experience rather than remaining isolated to a single suggestion type.

## Reformulating the code editing task

Knowing that a unified model was the north star, we realized this was an opportunity to reformulate the code editing problem, starting from the model output format. None of the three individual siloed tasks of predicting a suffix (completions), a window rewrite (NES), or a line number (cursor jump), which we had solved systematically and individually over time, were sufficiently expressive to represent the other tasks entirely. To combine these into a single model that could do all three, we needed a similarly unified output format—this gave rise to an elegant solution that we call the **“diff patch” output format**.

Consider a developer changing a function signature. The next useful action may be completing a new argument at the cursor. A moment later, it may be updating a call site below. After that, it may be adjusting validation logic elsewhere in the file. These are different interactions, but they are part of one editing task. A shared patch language lets the model reason about them as **sequential steps of the same problem**, which can be presented to the user with greater fluidity.

To fully understand the model output format changes, one can think of the three individual model tasks each as natural extensions of LLM next-token prediction capabilities—completions **continued unfinished code**, NES **rewrote a local edit window**, and long-distance edits **predicted the next line number to jump to**, accompanied by a call to NES. The 3-in-1 model puts these together to tackle all three at once while allowing the model to predict **multiple edits in one pass**.

To make this concrete, consider this example, where the symbol `<|cursor|>` represents the position of the user’s cursor.

```
def feed_penguin(penguin, fo<|cursor|>):
    penguin.meals += 1
    penguin.hungry = False
    penguin.last_ate_time = datetime.now()

    print(f"{penguin.name} ate {fish.type}")

penguin = Penguin("emperor")
feed_penguin(penguin, fish="silverfish")
```

The completions, NES, and long-distance edit models might predict the following in isolation to make the best edit for their corresponding tasks:

| Completions format | NES (Window Rewrite) format | Long-Distance Edit format |
| --- | --- | --- |
| `od` | ``` def feed_penguin(penguin, food):     penguin.meals += 1     penguin.hungry = False     penguin.last_ate_time = datetime.now()      print(f"{penguin.name} ate {food.type}") ``` | `5` (the next line number where fish should be replaced with food) |

The diff patch output format puts all of these elements together into a chain of edits that makes most sense at the cursor, farther away from the cursor, and even in a related file:

```
/path/to/penguins.py:0
-def feed_penguin(penguin, fish):
+def feed_penguin(penguin, food):
/path/to/penguins.py:5
-    print(f"{penguin.name} ate {fish.type}")
+    print(f"{penguin.name} ate {food.type}")
/path/to/penguins.py:8
-feed_penguin(penguin, fish="silverfish")
+feed_penguin(penguin, food="silverfish")
/path/to/antarctica.py:42
-feed_penguin(penguin2, fish="krill")
+feed_penguin(penguin2, food="krill")
```

Multi-patch responses also importantly allow us to **cache subsequent patches** (in the example above, the edits at lines 5, 8, and 42). Upon acceptance of preceding patches, they can appear with low latency, which results in a **fast tab-tab-tab code editing experience**.

This format expresses the three previous tasks in full and allows for flexibility in elegantly expressing future behaviors like cross-file edits (edits suggested in a related, but different file from the current one). More general formats such as tool calls are more abundant in pretraining and support actions beyond editing, but for code edits, we found that this format is most compact while remaining expressive.

In addition to reformulating the output task, we also took the opportunity to revisit the context given to the model as input. Originally, the input to the model included:

1. Recently viewed code snippets without line numbers
2. Current file content with line numbers
3. Recent edit history of the user in unified diff format
4. Area around the code to edit, a total of +/- 15 lines around the user’s cursor to pad the code to edit region
5. Code to edit, the rewrite window spanning 2 lines above and 5 lines below the user’s cursor

The unified output format enabled us to **simplify the input context as well**. Because the task scope now allowed for arbitrary edits anywhere in the file and was not limited to predicting the suffix at a specific location or rewriting a fixed window, we could remove the last two sections of the input—area around the code to edit and code to edit—and instead replace it with a single 3-line block with the line number and file content of the current cursor location. One important learning we carried forward from the previous models was that since the user’s cursor is constan
soheilpro11
🟧 echo.blog ⭐GitHub describes a single 3-in-1 inline-suggestions model using a diff-patch output format, with subsequent patches cached to provide fasterJulia Gong, Ben Liggett, and Ulugbek Abdullaev——

Interpretation history

Decision trace