2026-10-11 16:37 UTC

Microsoft's October 7 keynote and companion first-party blogs claim local LLM inference is now a first-class Windows path — Windows ML shipping experimental llama.cpp/GGUF support today, DeepSeek V4 Flash running locally in 60GB on RTX Spark, and GitHub Copilot gaining local models (MAI Code 1.1 Flash at ~70.8% SWE-Bench Verified on-device) with MXC-sandboxed tool execution by end of October — and on-schedule Copilot shipping plus real developer adoption of the Windows ML stack confirms local inference as mainstream on Windows, while slippage or quiet fade refutes it.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highlocal-inference windows-ml llama-cpp coding-agents agent-sandboxingMicrosoftGitHubNVIDIAPavan Davuluri

What is this?

Microsoft held a Windows and Surface event on October 7, 2026, accompanied by first-party devblogs from Windows ML, GitHub Copilot, and Windows Developer divisions. The coordinated announcements claim Windows ML now ships experimental llama.cpp/GGUF support with an OpenAI-compatible endpoint (live), Microsoft Execution Containers (MXC) reached GA as a cross-platform containment layer for agent tool execution, and GitHub Copilot will add MAI Code 1.1 Flash (~70.8% SWE-Bench Verified quantized on-device) with explicit local model selection and MXC sandboxing by end of October. Dated checkpoints include Nemotron >70B and HydraFusion local support on October 15. The web snippets confirm the October 7 event date and the Build 2026 (June) precursors (MAI models, MXC, RTX Spark Dev Box), but do not contain the October 7 keynote content itself — only the save-the-date page and June Build recaps are visible in search results.

Why it matters to Scott

Microsoft's October 7 stack ships the exact architecture Scott has been building and arguing for: MXC is a Windows-native ProcessContainer/MicroVM containment layer matching SiloOS and Runtime Containment patterns; Copilot's MAI Code 1.1 Flash (~70.8% SWE-Bench Verified on-device) with explicit local model picker and MXC sandboxing converges with his Agentic Coding loop (plan/act/reflect/recover), Code-First Architecture, and Test-First Agent Workflow; Windows ML adopting llama.cpp/GGUF with OpenAI-compatible endpoint validates his open-weights-on-Windows and hardware-aware local inference positions; MAI models flowing Foundry→Windows ML→Copilot confirms his Microsoft AI stack coherence thesis. This is a consequential vendor independently arriving at his load-bearing positions — dated-receipts opportunity.
ip:framework.siloosip:concept.runtime-containmentip:concept.architectural-containmentip:concept.capability-scope-separationip:concept.capability-tokensip:source.agentic-coding-plain-and-spicy-ebookip:framework.code-first-architectureip:concept.code-as-step-between-model-runsip:source.think-in-whole-stories-why-ai-coding-agents-write-better-code-when-they-see-the-complete-picture-ebookip:source.give-the-agent-a-workshop-ebookip:concept.test-first-agent-workflowip:concept.plan-mode-disciplinedev:project.all-in-one-softwaredev:project.beamdev:project.askdev:concept.hardware-aware-local-inferenceradar:abliterated-weights-agent-backdoorradar:aa-agentperf-local-benchmarkradar:ante-offline-coding-agentradar:adaptive-kv-cache-streamingradar:aegis-inline-ebpf-agent-containmentradar:antigravity-sdk-local-modelsradar:atlas-local-inference-engineradar:axera-ax8850-gguf-runtime
queries asked of Scott's wikis
  • local-inference strategy: Windows ML vs llama.cpp vs ONNX Runtime, model sovereignty on Windows
  • agent sandboxing: MXC / ProcessContainer / MicroVM threat model, capability grants, Windows vs Linux parity
  • coding agents: on-device SWE-Bench verified models, MAI Code family, Copilot local model picker UX
  • open-weights on Windows: GGUF/llama.cpp as first-class path, hardware acceleration (RTX Spark, NPU, DirectML)
  • developer adoption signals: Windows ML SDK uptake, Copilot local model opt-in rates, VS/VS Code integration depth
  • Microsoft AI stack coherence: Foundry / Azure AI / Windows ML / GitHub Copilot model sharing, MAI model provenance

Measured heat

now 4 pts/hpeak 53 pts/hcomments 2/hpeers p60momentum: steady2 platformsage 123h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-08 05:44⭐ origin directly observedMicrosoft Execution Containers: Policy-driven containment for AI agents
madspindel on hacker news
—
10-06 13:00first on blog (echo) · published · +-40.7hOriginal Microsoft announcement: "Today we are highlighting improvements across a few key open-source projects many of you already use, as w
Microsoft (Anastasiya Tarnouskaya, Michael Von Hippel, Tucker Burns — Microsoft Foundry on Windows Blog)
—
10-07 19:40first on hacker news · published · +-10.1hAI Development on Windows: From PyTorch and Llama.cpp to Windows ML
antimora
—
10-07 19:40amplified on hacker newshn.story.49997807
antimora
peak 3 · 0 comments · 1% of case engagement
10-07 19:59amplified on hacker newshn.story.49998043
choult
peak 2 · 0 comments · 0% of case engagement
10-07 20:54amplified on hacker newshn.story.49998660
lisajaloza
peak 2 · 1 comments · 1% of case engagement
10-08 05:44amplified on hacker newshn.story.50002144
madspindel
peak 5 · 0 comments · 1% of case engagement
10-08 07:06amplified on hacker newshn.story.50002660
pjmlp
peak 1 · 0 comments · 0% of case engagement
10-08 13:34amplified on hacker newshn.story.50005641
AllForAll
peak 2 · 0 comments · 0% of case engagement
2 more amplifiers in ainews.case_chain
10-07 20:21our radar first saw it · +-9.4hdiscovery anchor: hn.story.49997807—
pace: p77 vs 1247 stories at the 96h mark (now 123h old) — ahead of claude-code-remote-attribution-injection (1.0x), behind cloudflare-security-audit-skill (1.0x)

Evidence (9) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAI Development on Windows: From PyTorch and Llama.cpp to Windows ML
Retrieved article excerpt

Open article · Retrieved 2026-10-07T22:34:05.271448+00:00

Today we are highlighting improvements across a few key open-source projects many of you already use, as well as updates to Windows ML. To empower developers, we must support the broad range of tools for experimentation and exploration across inference and training, in addition to our production grade native inference stack.

Available today, Windows adds experimental llama.cpp support to Windows ML so you can run GGUF models locally through new task-specific APIs. Windows ML is the unified, high-performance local AI inferencing framework for Windows. Our experimental Windows-native Runtime API is now in preview for developers who want more control over how models run and compose.

These updates arrive with improvements across the open-source projects many of you already use, including PyTorch, and Triton. On a new generation of powerful Windows PCs powered by NVIDIA RTX Spark, like Surface Laptop Ultra, they give you more ways to run open-source models, build local agentic systems, and develop and optimize your own models on Windows.

## Building with the llama.cpp open-source community

We are very excited about advancements from across the open-source community – and nothing more so than the work happening across the **GGUF** and **llama.cpp** ecosystem. On this class of device, great support for the open-source models and frameworks developers already use matters most – and more of it is coming to Windows, natively and across Windows on Arm, than ever before.

We’re bringing **GGUF support** to **Windows ML**. The GGUF community is moving fast, unlocking new use cases with the latest open-source models and workloads, and llama.cpp is how many developers run them first. With this new, experimental llama.cpp integration, you can pull a brand-new GGUF model from Hugging Face and run it locally through the same Windows ML stack you already use, in just a few lines of code.

We’re also contributing directly to llama.cpp. Together with **NVIDIA** and the broader community, we’ve delivered significant performance improvements – through CUDA kernel optimization, kernel fusion, improved CPU–GPU scheduling, weight repacking, and CUDA graphs. That work also adds Eagle-3, MTP, and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures, and backend sampling.

**Explore:** [llama.cpp](https://github.com/ggml-org/llama.cpp)

## New to Windows ML?

For those of you who are new to **Windows ML**, it is the **unified, high-performance local AI inferencing framework for Windows**. You can use it to run your own custom local AI workflows across Windows PCs spanning GPUs, NPUs, and CPUs from AMD, Intel, NVIDIA, and Qualcomm. Local inference can help reduce latency, keep workload data on the device, and avoid per-token cloud inference charges. It’s designed to meet you where your models already live – whether you’re bringing something you trained and exported yourself or a popular open-source model from the community.

Windows ML also includes the [**Windows ML CLI**](https://aka.ms/winmlcli), a command-line tool and set of agent skills for getting your models ready to run with Windows ML. You can use it to convert, optimize, compile, and benchmark your models before you ship.

Many app developers are already using the Windows ML stack to enable local AI in their applications:

[App developers leveraging Microsoft Foundry on Windows to enable local AI in their applications](https://devblogs.microsoft.com/foundry-on-windows/wp-content/uploads/sites/94/2026/10/Fall_26_Logo_Soup.webp)

With this release, Windows ML adds an additional, experimental **Windows-native inferencing path** that runs both ONNX and GGUF models, making it easier to experiment with popular open-source models alongside the ONNX workflows you already use.

[Windows ML flow with Current ONNX Runtime Path](https://devblogs.microsoft.com/foundry-on-windows/wp-content/uploads/sites/94/2026/10/WinML_Flow_Diagram.webp)

## Experiment with open-source models and GGUF using Windows ML

So how do you run a GGUF model on Windows ML? Windows now includes task-specific APIs, starting with the Windows ML **Text Generation API**. It accepts language models in both **GGUF** and **ONNX** formats through one simplified surface. Windows ML automatically selects the right execution engine for your model – including llama.cpp for GGUF – so you can bring your own model, get results back quickly, and focus on building your application instead of the underlying plumbing.

For the quickest way to experiment, these APIs also expose an **OpenAI-compatible endpoint**. You can prototype against a local, on-device model using the same OpenAI SDK you already know – no new API to learn.

```
# Start the local Windows ML server with your GGUF model, then point the OpenAI SDK at it:
#   WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8080/v1",
    api_key="<access key printed on startup>",
)

stream = client.chat.completions.create(
    model="qwen2.5-0.5b",
    messages=[{"role": "user", "content": "What workloads can I run locally on the powerful NVIDIA RTX GPU in my Surface Laptop Ultra?"}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
```

This first release ships two task-specific APIs: the **Text Generation API**, which runs your own GGUF or ONNX language model, and the **Speech Recognition API**, which transcribes audio with an ONNX Whisper model. You can also use them together – transcribe voice input with the Speech Recognition API, then pass that text to the Text Generation API running a GGUF model. Just supply the models, and Windows ML handles the rest.

**Explore:** [Text Generation API docs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/text-generation) · [Speech Recognition API docs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/speech-recognition)

## The Windows ML Runtime API: a new Windows-native inferencing path

Underneath those higher-level surfaces is the **Windows ML Runtime API** – a new, experimental Windows-native inferencing API for Windows ML, built for performance, deep OS integration, and precise control when you want it. The familiar **ONNX Runtime APIs** stay fully supported for broad compatibility, while the Runtime APIs are where deeper, Windows-native optimizations will land over time – and the two ship side by side, so you can start with what you know and adopt the native path when you’re ready.

With the Runtime APIs you can:

- **Work with Windows-native data types.** Feed images, video frames, audio buffers, and text directly to models through efficient, zero-copy paths – instead of hand-writing preprocessing and format conversions.
- **Compose deterministic multi-model pipelines.** Chain multiple models into a single pipeline with explicit, per-stage device placement across CPU, GPU, and NPU. Execution is reproducible and predictable run-to-run – which matters as apps grow richer, like an encoder feeding a decoder.
- **Load and compile models ahead of time.** Turn a model into a ready-to-run artifact for faster startup, with your device and execution-policy choices applied consistently.

**Drop down to the Runtime API when you need finer control.** The higher-level Text Generation and Speech Recognition APIs are built on it, so you can start high-level and reach for the Runtime primitives whenever you need to.

[Windows ML New Flow with ONNX Runtime APIs and new Windows ML Runtime APIs](https://devblogs.microsoft.com/foundry-on-windows/wp-content/uploads/sites/94/2026/10/WinML_Flow_Diagram_New.webp)

**Explore:** [Windows ML Runtime API](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/overview) · [Runtime API samples](https://aka.ms/winml-runtime-samples)

## Open-source AI frameworks for local development on Windows

Beyond llama.cpp, Windows is gaining a wider open-source stack, from model libraries to the tooling that ties them together, with more of it running natively on Windows and Windows on Arm.

### PyTorch and Triton on Windows on Arm

**PyTorch** is the framework most developers reach for across AI development – from research and experimentation to building, training, and fine-tuning models. On Windows, it’s also a natural on-ramp to Windows ML: you can train or fine-tune your own models, including proprietary ones, and then bring them to Windows ML for local inference.

PyTorch now offers official **native Windows Arm64** CPU builds, and NVIDIA publishes CUDA-enabled Windows Arm64 packages for supported hardware. This gives model developers a native foundation for training, fine-tuning, and inference on Arm-based Windows AI systems. The Windows distribution of Triton brings **triton.jit**, **torch.compile**, and custom GPU kernels to supported Windows GPUs. Windows Arm64 compiler and release work extends that path to the optimized kernels used throughout modern AI frameworks.

**Explore:** [PyTorch Arm native builds](https://blogs.windows.com/windowsdeveloper/2025/04/23/pytorch-arm-native-builds-now-available-for-windows/) · [NVIDIA Windows Arm64 PyTorch packages](https://pypi.nvidia.com/nvtorch_oot/) · [Triton for Windows](https://github.com/triton-lang/triton-windows)

### Build with native PyTorch and Triton, then run with Windows ML

**PyTorch** and **Triton** help us complete the full model lifecycle path on Windows. In this example, PyTorch loads a real vision model**, torch.compile** uses Triton-generated GPU kernels on Windows, the original model is exported to the open **ONNX format** and prepared for deployment in an app.

#### Set up the native environment with PyTorch and Triton

```
pymanager install 3.14-arm64
pymanager exec -V:3.14-arm64 -m venv .venv
.\.venv\Scripts\Activate.ps1

python -m pip install --extra-index-url https://pypi.nvidia.com/nvtorch_oot `
    "torch==2.14.0+cu134" "torchvision==0.29.0+cu134"
python -m pip install "triton-windows==3.8.0.post29" `
    onnx onnxscript numpy
```

The benefit of Triton is visible when PyTorch Inductor combines a chain of GPU operations into a generated kernel. This small activation block creates several eager operations but can be fused by torch.compile for significant performance gains you can see comparing eager to triton:

```
import torch

def activation_block(x):
    return torch.nn.functional.silu(x * 1.5 + 0.25).square()

def time_ms(fn, x, iterations=100):
    for _ in range(20):
        fn(x)
    torch.cuda.synchronize()
    start = torch.cuda.Event(enable_timing=True)
    end = torch.cuda.Event(enable_timing=True)
    start.record()
    for _ in range(iterations):
        fn(x)
    end.record()
    torch.cuda.synchronize()
    return start.elapsed_time(end) / iterations

x = torch.randn(4096, 4096, device="cuda")
triton_block = torch.compile(activation_block, backend="inductor")
torch.testing.assert_close(
    triton_block(x), activation_block(x), rtol=1e-3, atol=1e-3
)

eager_ms = time_ms(activation_block, x)
triton_ms = time_ms(triton_block, x)
print(f"Eager:  {eager_ms:.3f} ms")
print(f"Triton: {triton_ms:.3f} ms")
print(f"Speedup: {eager_ms / triton_ms:.2f}x")
```

PyTorch remains the programming model while Inductor generates specialized Triton GPU code to be used while working on the model efficiently.

#### Export the portable ONNX model

The model graph—not the generated kernel—is exported for deployment.

```
import torch
from torchvision.models import resnet18

model = resnet18(weights=None).eval()
image = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model,
    (image,),
    "resnet18.onnx",
    dynamo=True,
    external_data=False,
    input_names=["image"],
    output_names=["scores"],
)
```

#### Build once, benchmark, and ship

The Windows ML CLI exposes analyze, optimize, quantize, and compile to developer
antimora30
🟧 hnLocal models and sandboxed tools coming soon to GitHub Copilot
Retrieved article excerpt

Open article · Retrieved 2026-10-07T22:34:05.853909+00:00

[Deep Dive](https://commandline.microsoft.com/?type=Deep%20Dive)

Share

[in](https://commandline.microsoft.com/local-models-sandboxed-tools-github-windows/ "Share on LinkedIn")
[r](https://commandline.microsoft.com/local-models-sandboxed-tools-github-windows/ "Share on Reddit")
[x](https://commandline.microsoft.com/local-models-sandboxed-tools-github-windows/ "Share on X")
[f](https://commandline.microsoft.com/local-models-sandboxed-tools-github-windows/ "Share on Facebook")

Bringing local models and sandboxed tools to Windows and GitHub Copilot

# Bringing local models and sandboxed tools to Windows and GitHub Copilot

Coming soon, GitHub Copilot will determine when a task is best handled by on-device intelligence and when it should leverage cloud-scale models.

By [Patrick Nikoletich](https://commandline.microsoft.com/author/patrick-nikoletich)Distinguished Product Manager, GitHub and [Stuart Schaefer](https://commandline.microsoft.com/author/stuart-schaefer)Partner Architect, Windows Platform, Microsoft

2026.10.07

When leveraging agents, developers need both choice and control. They need technologies that offer clear boundaries for their agents and make it easy to choose the model with the right speed, performance, and cost profile for each task. That’s why GitHub offers frontier models from major model providers, as well as options like [Project HydraFusion](https://github.blog/ai-and-ml/github-copilot/project-hydrafusion-frontier-quality-via-multi-model-orchestration/?utm_source=blog-cta-project-hydrafusion-article&utm_medium=blog&utm_campaign=oct-7-event-oct-2026), an orchestrator choosing one or multiple models for each task while balancing performance, cost, and latency. It’s also why Windows has developed [Microsoft Execution Containers (MXC)](https://aka.ms/WindowsDeveloperMXC) to help secure interactive and non-interactive agentic coding sessions.

Coming by the end of the month, GitHub Copilot will determine when a task is best handled by on-device intelligence and when it should leverage cloud-scale models. Rather than forcing developers to manage infrastructure decisions themselves, GitHub Copilot automatically coordinates local and cloud inference behind the scenes. For NVIDIA RTX Spark Windows PCs like Surface Laptop Ultra, that means we are enabling local coding in GitHub Copilot with powerful local inference models and hardware capable of delivering a great experience at the edge.

The result is poised to be the next step in the HydraFusion vision: intelligent orchestration that spans not just multiple models, but multiple compute environments including the edge. GitHub Copilot can run commands in these environments with controlled access to files, networks, system capabilities, and credentials. Developers can automate with confidence and security in mind.

## Why memory matters for a local coding agent

Performance readout for a Surface Laptop Ultra with an NVIDIA RTX Spark GPU.

Local inference starts with a memory budget. Surface Laptop Ultra is built around NVIDIA RTX Spark, with up to 128 GB of unified memory and up to 1 petaflop of AI compute.

With a discrete GPU, dedicated video memory is an important constraint: moving model data between system memory and the GPU can add overhead. Unified memory gives the CPU and GPU access to a shared physical pool. That makes more capacity available to the workload, but it doesn’t make all of it available to model weights.

The operating system, your applications, and the inference runtime need memory, too. So does the key-value cache, which stores attention state for tokens the model has already processed. As an agent reads files and receives tool results, its context can grow, increasing memory use and the work needed to process the next request.

A breakdown of the CPU and GPU's shared memory budget.

Model weights are only part of the memory budget.

Keeping a model loaded between requests can avoid repeated loading work. However, it doesn’t guarantee constant response time; context length, memory pressure, and the rest of the workload still matter. That’s why the useful question isn’t just whether a model fits, but how it behaves over a complete coding task.

## Introducing MAI Code 1.1 Flash for local coding

To bring this local-development experience to life, Microsoft AI developed a local version of MAI Code 1.1 Flash, a coding-optimized mixture-of-experts model with a 137 billion total and 6.8 billion active parameters. The on-device work applies quantization and speculative decoding to reduce the model footprint and improve end-to-end responsiveness while preserving the task completion and tool-use quality that matter in an agent loop.

Quantization reduces the precision used to represent model weights and activations, lowering memory requirements. Because code doesn’t degrade gracefully, the release evaluation must measure coding-task success as well as footprint: a single incorrect token can produce a syntax error, wrong identifier, malformed tool call, or broken diff.

Speculative decoding trades additional working memory for higher decode throughput and lower end-to-end latency. A drafter proposes candidate token blocks and the target model verifies them.

For a coding agent, the important tradeoff is whether the smaller model can still complete the same tasks. A smaller footprint is useful only if changes in code quality and tool use are understood.

With our first shipping version of MAI Code 1.1 Flash on Surface Laptop Ultra we achieve the following performance at different context lengths, with peak memory usage of 75.5GB at 256k context. At 64k and 128k context, prompt-processing throughput reaches 923.5 and 769.8 tokens per second, respectively.

A chart showing decode throughput of MAI Code 1.1 Flash, based on prompt length.

The quantized version of MAI Code 1.1 Flash we use on device retains capability impressively compared to the Bfloat16 cloud variant, coming in at 53GB, an 80% reduction in size.

| **Benchmark** | **Dataset size** | **MAI Code 1.1 Flash** | **GPT OSS 120B\*** | **MAI Code 1.1 Flash Quantized on Device** |
| --- | --- | --- | --- | --- |
| SWE-Bench Verified | 500 | 72.6% | 32.0% | 70.80% |
| Terminal-Bench 2.1 | 89 | 62.9% | 23.6% | 66.29% |

\*GPT OSS version for comparison was [Unsloth’s GPT-OSS-120B GGUF](https://huggingface.co/unsloth/gpt-oss-120b-GGUF).  
  
Tested October 5, 2026 using MAI Code 1.1 Flash (mixed-precision quantization, approximately 3.3 bits per weight) with DFlash2 sliding-window speculative decoding and a Windows ARM64 llama.cpp CUDA runtime. Results reflect decode throughput for a synthetic code-generation workload; actual results may vary by device, configuration, and other factors.

## Two ways to use local models in GitHub Copilot

GitHub Copilot is adding two ways to use local models across the GitHub Copilot CLI, Copilot app, and VS Code. Developers can let Copilot’s intelligent Auto orchestration choose when to use local or cloud inference, or they can explicitly select a local model for workflows that require direct control.

Developers can let Copilot’s intelligent Auto orchestration choose when to use local or cloud inference, or they can explicitly select a local model for workflows that require direct control.

Model selection, inference, and tool execution have different boundaries; local inference does not make the session offline.

With Auto, developers do not need to decide where each task should be run. Across a multi-turn session, Copilot can consider task context and cache state as it routes work between local and cloud models, preserving useful, cached work as the session evolves.

This orchestrated experience complements direct model selection, giving developers a choice between letting Copilot optimize model placement and choosing a specific local model themselves.

Explicit local-model selection supports workflows that need a specific provider, model, or endpoint. Developers can select MAI Code 1.1 Flash through the Windows ML provider or connect GitHub Copilot to OpenAI-compatible local endpoints and choose from the models those endpoints expose.

## How sandboxes help secure tool execution

An agent’s shell commands normally inherit the access of the account running them. Moving inference onto the device doesn’t change that. Sandboxing applies a policy to the processes and local services the agent launches, controlling access to files, networks, credentials, system capabilities, and execution paths regardless of which model requested the work.

Sandboxing applies a policy to the processes and local services the agent launches, controlling access to files, networks, credentials, system capabilities, and execution paths regardless of which model requested the work. GitHub Copilot uses Microsoft Execution Containers, or MXC, an open-source library from the Windows team that translates policy into native operating-system controls.

GitHub Copilot uses [Microsoft Execution Containers](https://aka.ms/WindowsDeveloperMXC), or MXC, an open-source library from the Windows team that translates policy into native operating-system controls. On Windows, GitHub Copilot uses the BaseContainer tier of the ProcessContainer backend. On macOS, it uses Seatbelt. On Linux, it uses [bubblewrap](https://github.com/containers/bubblewrap). These local backends don’t require a separate virtual machine or container image, but we plan to make them options available through MXC in the future.

When sandboxing is enabled, shell commands and, by default, local Model Context Protocol servers and language servers run inside the process boundary. Built-in file tools run inside GitHub Copilot itself: the agent harness checks their requests against the effective policy, but those checks aren’t OS-enforced child-process isolation. Remote MCP servers are also outside the local process sandbox; when MCP sandbox controls apply, GitHub Copilot checks their connection policy in process.

Opening the GitHub Copilot CLI and running `/sandbox` slash command allows you to configure your settings at any time.

## A real-world example: Daily repository dashboard

To show how these technologies work together, let’s use a real-world example.

Consider a job that reads local repositories, runs their tests in working copies, and writes one HTML report each morning. The source repositories should remain read-only, and the test processes shouldn’t access the network. Model selection is independent: the same job can use a configured local or cloud model.

### Enable sandboxing for your project

Open the settings dialog by clicking the gear icon in the GitHub Copilot app and selecting your project in the left menu. For this example, we’ll be using the `copilot-sdk` repo that hosts our opensource GitHub Copilot Runtime/SDK project.

Enabling the `Sandbox new sessions` toggle turns on sandboxing by default whenever you work within that project. By default, it’s current working directly is read/write while the rest of the system remains largely read-only or inaccessible to an agent.

### Create a new automation

Select the` Automations` section on the left navigation and click the `Start automation` button to open the dialog that allows you to configure a new automation.

After setting a clear title and a trigger time of 9AM daily, paste in basic instructions to guide the agent through the creation of a daily dashboard for triaging.

Copy
`Create today’s public triage dashboard for github/copilot-sdk using only GitHub issue and PR metadata (no local repos, code, tests, or off-repo links), summarizing open/closed/merged items, recently updated work, stale items, labels, authors, assignees, and age; generate .\dashboard\index.html with inline CSS and SVG, append today’s results to .\history.json for up to seven dates, and finish with exactly three lines: report path, failures, and skipped or unavailable work.`

Selecting the new MAI Code 1.1 Flash local model and `copilot-s
lisajaloza21
🟧 hnMS announces DeepSeek V4 Flash and Nemotron locally, llama.cpp comes to Win ML
Retrieved article excerpt

Open article · Retrieved 2026-10-07T22:34:14.635553+00:00

## Accessibility

×

Theme

Light theme
Dark theme
High contrast


Text size

A−
100%
A+


Line spacing

Normal
Wide


Paragraph spacing

Normal
Wide

Underline all links

Reduce animations

Readable font

Larger cursor

Reset all
OK

news

Pasquale Pillitteri ingegnere informatico sviluppo software Palermo

Pasquale Pillitteri

Accessibility options

EN 

[IT](https://pasqualepillitteri.it/news/21456/deepseek-v4-flash-nemotron-locale-windows-rtx-spark) 
[EN](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml) 
[FR](https://pasqualepillitteri.it/fr/news/21458/windows-deepseek-v4-flash-nemotron-local-llama-cpp) 
[ES](https://pasqualepillitteri.it/es/news/21459/windows-anuncia-deepseek-v4-flash-nemotron-local) 
[DE](https://pasqualepillitteri.it/de/news/21460/windows-deepseek-v4-flash-nemotron-lokal-llama-cpp) 
[TR](https://pasqualepillitteri.it/tr/news/21461/deepseek-v4-flash-nemotron-yerel-windows-ml) 
[RU](https://pasqualepillitteri.it/ru/news/21462/windows-deepseek-v4-flash-nemotron-llama-cpp) 
[ZH](https://pasqualepillitteri.it/zh/news/21463/windows-xuanbu-deepseek-v4-flash-nemotron-bendi-yunxing) 
[PT](https://pasqualepillitteri.it/pt/news/21464/windows-anuncia-deepseek-v4-flash-nemotron-local) 
[JA](https://pasqualepillitteri.it/ja/news/21465/windows-deepseek-v4-flash-nemotron-locale-llama-cpp-windows-ml) 
[ID](https://pasqualepillitteri.it/id/news/21466/windows-deepseek-v4-flash-nemotron-lokal-llama-cpp) 
[PL](https://pasqualepillitteri.it/pl/news/21467/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml) 
[NL](https://pasqualepillitteri.it/nl/news/21468/windows-deepseek-v4-flash-nemotron-lokaal) 
[HI](https://pasqualepillitteri.it/hi/news/21469/windows-deepseek-v4-flash-nemotron-lokal-llama-cpp)

Certificazione Innovation Manager UNI 11814 Palermo  
UNI 11814:2021

- [Advertising](https://pasqualepillitteri.it/en/banners "Advertise on pasqualepillitteri.it")
- [Privacy](https://pasqualepillitteri.it/en/privacy "Privacy Policy Pasquale Pillitteri")
- [Cookies](https://pasqualepillitteri.it/en/cookies "Cookie Policy")
- [Accessibility](https://pasqualepillitteri.it/en/accessibility "Accessibility")
- [Terms](https://pasqualepillitteri.it/en/terms "Terms & Conditions")
- [Standards](https://pasqualepillitteri.it/en/editorial-standards)
- [Corrections](https://pasqualepillitteri.it/en/corrections)
- [Ethics](https://pasqualepillitteri.it/en/ethics)

© 2026

[P. Pillitteri](https://pasqualepillitteri.it/)

[FREE **Subscribe to the Saturday newsletter** · 3.4k readers worldwide, every Saturday Subscribe →](https://pasqualepillitteri.it/en/newsletter)[Tech Gear **The right hardware for AI** · Mini PCs and GPUs for local LLMs, edge boards and hand-picked books. Browse the board →](https://pasqualepillitteri.it/en/tech-gear#ai)×


# Windows announces DeepSeek V4 Flash and Nemotron locally, llama.cpp comes to Windows ML

##### At the Windows keynote, Microsoft brings DeepSeek V4 Flash and Nemotron to RTX Spark and puts llama.cpp in Windows ML. What we know about V4 Flash so far.

[Most Read 
Most Read](https://pasqualepillitteri.it/en/news/576/claude-code-skills-design-uiux-guide "Most Read")
[AI News & Trends 
AI News & Trends](https://pasqualepillitteri.it/en?cat=77 "AI News & Trends")
[Gaming 
Gaming](https://pasqualepillitteri.it/en?cat=414 "Gaming")
[Claude Code & Anthropic 
Claude Code & Anthropic](https://pasqualepillitteri.it/en?cat=62 "Claude Code & Anthropic")
[Google AI & Gemini 
Google AI & Gemini](https://pasqualepillitteri.it/en?cat=67 "Google AI & Gemini")
[AI Tools & Reviews 
AI Tools & Reviews](https://pasqualepillitteri.it/en?cat=72 "AI Tools & Reviews")
[Apple 
Apple](https://pasqualepillitteri.it/en?cat=157 "Apple")
[Cybersecurity 
Cybersecurity](https://pasqualepillitteri.it/en?cat=21 "Cybersecurity")
[Benchmarks & Comparisons 
Benchmarks & Comparisons](https://pasqualepillitteri.it/en?cat=103 "Benchmarks & Comparisons")
[Automotive Tech 
Automotive Tech](https://pasqualepillitteri.it/en?cat=386 "Automotive Tech")
[Chess & AI 
Chess & AI](https://pasqualepillitteri.it/en?cat=456 "Chess & AI")
[Apple 
Apple](https://pasqualepillitteri.it/en?cat=158 "Apple")
[Agent Infrastructure 
Agent Infrastructure](https://pasqualepillitteri.it/en?cat=108 "Agent Infrastructure")
[Reports & Analysis 
Reports & Analysis](https://pasqualepillitteri.it/en?cat=57 "Reports & Analysis")
[OpenAI & ChatGPT 
OpenAI & ChatGPT](https://pasqualepillitteri.it/en?cat=147 "OpenAI & ChatGPT")
[Finance 
Finance](https://pasqualepillitteri.it/en?cat=87 "Finance")
[AI for Professionals 
AI for Professionals](https://pasqualepillitteri.it/en?cat=83 "AI for Professionals")
[AI News & Trends 
AI News & Trends](https://pasqualepillitteri.it/en?cat=78 "AI News & Trends")
[Software Development 
Software Development](https://pasqualepillitteri.it/en?cat=23 "Software Development")
[AI for Professionals 
AI for Professionals](https://pasqualepillitteri.it/en?cat=82 "AI for Professionals")
[Cybersecurity 
Cybersecurity](https://pasqualepillitteri.it/en?cat=184 "Cybersecurity")
[Guides & Tutorials 
Guides & Tutorials](https://pasqualepillitteri.it/en?cat=30 "Guides & Tutorials")




Pasquale Pillitteri

[Pasquale Pillitteri](https://pasqualepillitteri.it/en/author)

07/10/2026
 8 min read

[AI News & Trends](https://pasqualepillitteri.it/en/news/cat/77)
[ai tools](https://pasqualepillitteri.it/en/news/tag/ai+tools)
[ai news](https://pasqualepillitteri.it/en/news/tag/ai+news)
[deepseek v4 flash](https://pasqualepillitteri.it/en/news/tag/deepseek+v4+flash)
[v4 flash](https://pasqualepillitteri.it/en/news/tag/v4+flash)
[nemotron](https://pasqualepillitteri.it/en/news/tag/nemotron)
[rtx spark](https://pasqualepillitteri.it/en/news/tag/rtx+spark)
[windows ml](https://pasqualepillitteri.it/en/news/tag/windows+ml)
[llama.cpp](https://pasqualepillitteri.it/en/news/tag/llama.cpp)
[local ai models](https://pasqualepillitteri.it/en/news/tag/local+ai+models)

Table of Contents

- [1.What Microsoft said about local models](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-1)
- [2.Windows ML and llama.cpp, what changes for developers](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-2)
- [3.The models mentioned, one by one](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-3)
- [4.How much memory you need on RTX Spark](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-4)
- [5.What is not confirmed](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-5)
- [6.Frequently asked questions](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-6)
- [7.Conclusions](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#toc-7)
- [8.Rate this article](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#rating-section)
- [9.Related Articles](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#related-section)
- [10.Looking for a Software Engineer?](https://pasqualepillitteri.it/en/news/21457/windows-deepseek-v4-flash-nemotron-llama-cpp-windows-ml#contact-section)

Advertisement

Microsoft announced at the October 7 Windows keynote in San Francisco (10 a.m. PT, 1 p.m. ET) that **DeepSeek V4 Flash** will be able to run locally on Windows, in 60 GB of memory. V4 Flash, a 284-billion-parameter model, is quantized to 1.6 bits, while NVIDIA will launch a new Nemotron of more than 70 billion parameters on October 15 and llama.cpp lands in Windows ML. The figures are read off the slides shown on stage, and the official Windows Blog and NVIDIA posts are not online yet.

Local AI models on Windows and RTX Spark, DeepSeek V4 Flash and Nemotron announced by Microsoft


The picture of local models that emerged from the October 7, 2026 Windows and Surface keynote.

We had already covered the rest of the event, with the Surface Laptop Ultra and the new PCs, in the [preview on the eve of the event about RTX Spark and the new Surface](https://pasqualepillitteri.it/en/news/16603/microsoft-nvidia-october-7-event-rtx-spark) and in the piece on the [Surface Laptop Ultra](https://pasqualepillitteri.it/en/news/3914/surface-laptop-ultra-nvidia-rtx-spark-microsoft). This article looks at one thing only, namely which AI models run on the PC instead of in the cloud, with how much memory, and what you need to use them.

Advertisement[This space is for sale](https://pasqualepillitteri.it/en/banners) 

Update from the evening of October 7, [the Surface Laptop Ultra costs from $2,599 and ships October 16 with preorders open](https://pasqualepillitteri.it/en/news/21471/surface-laptop-ultra-price-release-date-october-16), while a euro price is still missing.

## What Microsoft said about local models

Pavan Davuluri, head of Windows and Microsoft devices, did the talking on models. According to the live blogs, the common thread was so-called hybrid intelligence, the idea that part of the work runs on the PC and part in the cloud, while the cost of API tokens keeps climbing. Against that backdrop, four model announcements followed.

The first concerns **DeepSeek V4 Flash**. The slide presents it as "Announcing", with 284 billion parameters quantized to 1.6 bits and "near-frontier intelligence in 60GB of memory", meaning intelligence close to that of frontier models in 60 GB. The second is Nemotron, NVIDIA's family of open models, with a new local model of more than 70 billion parameters expected on October 15. The third is llama.cpp, announced inside Windows ML. The fourth is [HydraFusion](https://github.blog/2026/09/04/project-hydrafusion-frontier-quality-via-multi-model-orchestration/), the GitHub system that picks which model to use for each task (since September it has been a research preview inside Copilot CLI) and that from October 15, the slide says, also supports local models on Windows.

"DeepSeek V4 Flash" slide shown on stage, 284 billion parameters quantized to 1.6 bits and "near-frontier intelligence in 60GB of memory".

"DeepSeek V4 Flash" slide shown on stage, 284 billion parameters quantized to 1.6 bits and "near-frontier intelligence in 60GB of memory".  
Source: official live stream on the Windows YouTube channel ([YouTube](https://www.youtube.com/watch?v=ilmBGeGldrI))

Advertisement

HydraFusion slide, available from October 15 with support for local models on Windows.

HydraFusion slide, available from October 15 with support for local models on Windows.  
Source: official live stream on the Windows YouTube channel ([YouTube](https://www.youtube.com/watch?v=ilmBGeGldrI))

A note on the name. [Engadget](https://www.engadget.com/2279642/microsoft-windows-surface-event-2026-live-blog-nvidia-rtx-spark-laptop-ultra/) writes "DeepSeek V5 Flash", but the slide says V4 Flash, and DeepSeek has never announced a V5 (only unsourced rumors exist). The live blogs from [BGR](https://www.bgr.com/2279488/windows-surface-event-october-2026-liveblog-updates/) and [SlashGear](https://www.slashgear.com/2279963/microsoft-surface-windows-event-october-2026-liveblog/), by contrast, report V4 Flash, as does this whole article.

Advertisement[This space is for sale](https://pasqualepillitteri.it/en/banners) 

## Windows ML and llama.cpp, what changes for developers

Windows ML is the layer Windows uses to run AI models on the computer. According to [Microsoft's documentation](https://learn.microsoft.com/windows/ai/new-windows-ml/overview), it sits on top of ONNX Runtime (the open-source engine for model inference), taps NPUs, GPUs or CPUs through "execution providers" that Windows installs and updates on its own via Windows Update, and accepts models conve
choult20
🟧 echo.blogOriginal Microsoft announcement: "Today we are highlighting improvements across a few key open-source projects many of you already use, as wMicrosoft (Anastasiya Tarnouskaya, Michael Von Hippel, Tucker Burns — Microsoft Foundry on Windows Blog)——
🟧 hn ⭐Microsoft Execution Containers: Policy-driven containment for AI agents
Retrieved article excerpt

Open article · Retrieved 2026-10-08T07:01:55.196571+00:00

# Microsoft Execution Containers: Policy-driven containment for AI agents

Written By

- [Logan Iyer, Corporate Vice President, Windows Platform + Developer](https://blogs.windows.com/windowsdeveloper/author/logan-iyer-corporate-vice-president-windows-platform-developer/)


published

October 7, 2026

Agents are unlocking enormous productivity gains for customers, but their ability to work across files, networks, and applications can introduce new security risks. This can leave customers feeling like they only have two choices: give agents unrestricted access and hope nothing goes wrong or block them and lose the productivity benefits they provide. Neither option is acceptable.

That’s why we’re building Windows platform capabilities to help run and manage agents more securely, starting with:

- **Containment**: limiting what an agent can access and do.
- **Identity**: distinguishing an agent’s activity from a person.
- **Manageability**: giving organizations the tools to govern access and monitor agent activity.

[**Microsoft Execution Containers (MXC)**](https://blogs.windows.com/windowsdeveloper/2026/06/02/build-2026-furthering-windows-as-the-trusted-platform-for-development/), now generally available, provides the containment layer. Developers and IT administrators define the resources, like files and network destinations an agent can use and MXC uses the appropriate container to enforce those policies at runtime.  

Windows will also soon enable Microsoft Entra to help distinguish agent activity from user activity, ensuring users can remain productive even when agent access needs restrictions and extend Microsoft Agent 365 controls to local agents on-device, enabling IT teams to manage MXC containers, apply policies, and monitor agent activity.

### Why agents need a managed execution boundary

An agent cannot be its own security authority. It must run within a boundary defined by the developer or organization and enforced independently of the agent itself.

Consider a coding agent asked to update a website. The agent needs read and write access to the website repository and access to the development tools required to build and test the change. It may need to read production server configuration to understand how the application is deployed, but it should not be able to modify that configuration.

Without a managed execution boundary, the agent may decide that changing the server configuration is the fastest way to complete the task and potentially break the production site. The action can be reasonable from the agent’s perspective and still exceed the authority the developer intended to grant.

Containment creates that managed boundary: the agent can read and write the repository, read the server configuration, but is not authorized to access anything else unless access has been granted. If the agent attempts to modify the server configuration, the containment environment is designed to prevent the operation regardless of what the model, generated code, plugin, or tool decides to do.

### What is Microsoft Execution Containers (MXC)?

[Microsoft Execution Containers (MXC)](https://github.com/microsoft/mxc/tree/main) is a policy-driven execution layer for untrusted code or dynamically generated workloads. In agentic scenarios, developers can use MXC to contain model-generated output, plugins, tools, agent harness, or the entire agent. This limits what the agent workload can access by containing the scope of the impact if something goes wrong.  

Developers declare the resources a workload requires, such as files and network destinations, and MXC enforces the resulting boundary with the appropriate container. The policy remains outside the agent workload’s control, so the agent or generated code cannot grant itself additional access.

MXC also separates workload requirements from platform-specific containment details, greatly simplifying the developer experience. Developers integrate with a unified JSON configuration schema and multi-language SDK, while MXC maps the requested controls to the selected backends on Windows, macOS, or Linux.

Developers can apply the same containment model wherever an agent needs to run – from a local device to the cloud. With Windows 365 support for MXC now generally available, developers can run agents alongside their existing work on Cloud PCs, using the right isolation model to help keep agent execution separate and secure.

### Choose the containment level that fits the workload

Different workloads require different levels of isolation. A coding agent working in a repository may prioritize low latency and responsiveness, while an agent processing sensitive data or running untrusted code may require stronger isolation. MXC provides a spectrum of containment options so developers and organizations can select the right level of isolation for each workload.

Only Windows supports a session container, which runs an agent on the user’s device in a separate, OS-isolated session with its own local agent identity and isolated desktop, clipboard, UI, and input boundaries.

|  |  |  |  |
| --- | --- | --- | --- |
| **Backend** | **Availability** | **Best suited for** | **Important characteristics** |
| **Process container** | Windows 11, macOS, and Linux | Lightweight containment for responsive workloads, including model-generated code and tool execution. | Uses the platform-appropriate process sandbox, including AppContainer on Windows, Seatbelt on macOS, and Bubblewrap on Linux. |
| **Session container** | Windows 11 only | Long-running agents and automation that need a desktop or stronger separation from the interactive user. | Runs under a distinct Windows account and session, separating the agent’s desktop, clipboard, UI, input, and active session from the user. |
| **WSL container (WSLc)** | Windows 11 only | Linux-first agent toolchains and workloads that depend on the Linux package and development ecosystem. | Provides a Linux execution environment through WSL. |
| **MicroVM** | Windows 11 and Linux, experimental | Higher-risk workloads that benefit from a hardware-backed virtualized boundary. | Provides hardware-enforced isolation and full Linux workload compatibility. |

Each containment backend has distinct security properties and workloads should be evaluated for fit with containment backend.

### How MXC policy defines workload boundaries

MXC helps users and organizations delegate more work to agents while limiting them to the resources required for each task. Instead of giving an agent the full authority of the signed-in user, developers and IT can define an OS-enforced boundary around the agent’s workload. 

For the website scenario, an MXC policy could give a coding agent read and write access to its local source-code repository and access to required development tools such as Git, while preventing access to personal locations such as the user’s Documents folder. The policy could also block inbound and outbound network connections and access to the interactive desktop.  

|  |  |
| --- | --- |
| **Policy area** | **What it controls** |
| **Containment** | The isolation environment in which the workload runs, such as a process or session container. |
| **Process** | The command, arguments, working directory, environment, and other settings used to start the workload. |
| **File system** | Locations the workload can modify, locations it can read without changing, and locations it cannot access. |
| **Network** | Inbound and outbound connectivity, including whether the workload can connect to services through the host’s loopback interface. |
| **User interface** | Whether the workload can access or interact with the desktop and related UI resources. |

Adding MXC support is straightforward for agent developers – you can easily use your favorite coding agent to integrate the MXC SDK and draft an initial workload policy, then review, test, and refine the resulting controls.

### How workload and organizational policy work together

Agent developers declare the resources their agentic workloads may need. Organizations can also apply additional constraints through management policy, like with Microsoft Intune management policy. This allows the same agent to operate within different enterprise boundaries without requiring the agent developer to encode the organization’s security posture into the application.

Containment becomes more valuable when organizations can apply it consistently. Intune policy will soon be available to manage MXC process containers used by MXC-integrated agents on Windows 11. These policies will offer IT administrators control over how Windows evaluates container creation requests made by agents, and the resource boundaries that are enforced by those containers.

Agent developers should design their workloads to operate within boundaries that may be more restrictive than the default configuration. If organizational policy blocks a resource, the agent should explain that the task could not be completed within the available permissions, request an appropriate user or administrator action when supported, or choose a safe alternative. It should not silently fail.

### Observe and refine policy before enforcing it

Writing a least privilege policy can be difficult when you do not yet know every resource an agent workload requires. Only on Windows, MXC process containers can produce an agent activity report showing which resources a workload attempted to use to help craft a least-privilege policy.

MXC supports three operating modes:  

|  |  |  |  |
| --- | --- | --- | --- |
| **Mode** | **Ungranted access** | **Activity Report** | **Intended use** |
| **Enforcement** | Blocked | No | Run the workload with its production policy |
| **Learning** | Blocked and recorded | Yes | Diagnose failures and verify that the policy grants only required access |
| **Permissive** | Allowed and recorded | Yes | Observe agent activity without enforcing policy |

In Enforcement mode, MXC applies the policy without activity report. Granted operations proceed, and operations outside the boundary are restricted.

In Learning mode, MXC continues to enforce the configured boundary. An operation that has not been granted is blocked and recorded in a JSON activity report. This lets developers or IT reproduce containment failures and understand which resources the workload attempted to access.

In Permissive mode, MXC records access that the policy would have denied but allows the operation to continue. This is useful during policy authoring because the workload can be completed while developers or IT collect evidence about the resources it uses. Permissive mode does not bypass other applicable operating system or organizational restrictions.

When you bring an agent into an MXC container for the first time, the policy may block capabilities that the workload legitimately needs. These denials reveal where the boundary needs adjustment, helping you grant the required access without unnecessarily expanding the agent’s reach.

### Agent identity & attribution

Containment answers what an agent may do. Identity helps determine which agent performed an action. Coming soon, Windows will allow Microsoft Entra to distinguish agent activity from user activity in Microsoft Agent 365.

This separation will allow security teams to evaluate an agent’s behavior and risk independently from the person using the device. If an agent becomes compromised or violates policy, security controls will target the agent’s access to protected resources without blocking the employee’s access. As organizations deploy more agents, one misbehaving agent does not have to interrupt the employee’s access or productivity.

Agent-level attribution also improves investigation and governance. Administrators will be able to use Agent 365 to understand which agents are running, associate activity with a specific agent, investigate risky behavior, and apply policy to individual a
madspindel50
🟧 hnAI Development on Windows: From PyTorch and Llama.cpp to Windows MLpjmlp10
🟧 hnMicrosoft is sandboxing AI agents at the OS levelAllForAll20
🟧 hnMxc: Microsoft Execution Containers version 1.0.0smokel21460
🟧 hnMXC - a sandboxed code execution systemnreece15273

Interpretation history

Decision trace