2026-10-11 16:37 UTC

Mark Russinovich and coauthors claim weaker unaligned orchestrators recover otherwise unavailable harmful capabilities by composing individually permitted consultations with aligned frontier models, exposing a safety gap beyond single-interaction refusal controls.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumagentic-security model-evaluation capability-controlMark RussinovichBlake BullwinkelGiorgio SeveriCristian OvadiucAhmed Salem

What is this?

The case describes “Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs,” attributed to Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, and Ahmed Salem. Its reported claim is that weaker unaligned orchestrators can assemble harmful capabilities from individually permitted consultations with aligned models. The supplied web snippets do not establish this paper or its reported cybersecurity and CBRN results: they instead describe separate GRP-Obliteration research on removing safety alignment through training with harmful prompts and judge-model scores. A Microsoft Security Blog snippet identifies Russinovich as Microsoft Azure CTO and Severi as a senior AI safety researcher, but does not substantiate the case’s consultation-based mechanism.

Why it matters to Scott

If substantiated, the reported mechanism would converge with Scott’s Guardrail Illusion position and extend his Model-Plus-Harness Benchmark Unit claim into safety evaluation: his model-to-model delegation and SiloOS designs would need tests of cumulative capability across permitted consultations, not just individual calls. The supplied snippets establish neither this paper nor its results, so this is a provisional evaluation lead—not yet a publishing receipt; the radar’s MCP tool-sequence bypass case tracks a related mechanism, not this development.
ip:concept.guardrail-illusionip:concept.model-plus-harness-benchmark-unitdev:concept.model-to-model-delegationdev:project.silo-osradar:mcp-tool-sequence-guardrail-bypass
queries asked of Scott's wikis
  • agent orchestration cross-call safety composition
  • capability control versus single-response refusal
  • local models frontier model delegation trust boundaries
  • agent harness end-to-end safety evaluation
  • task decomposition cumulative intent tracking

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 674h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-13 14:00⭐ origin echo-reconstructedThe authors report consultation-aided uplift on cybersecurity benchmarks and harmful CBRN requests when local orchestrators decompose tasks
Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, and Ahmed Salem on paper (echo) · attributed from hn.story.49720533
—
09-16 00:07first on hacker news · published · +58.1hDivide, Consult, Conquer: Capability Laundering Through Aligned LLMs
sbulaev
—
09-16 00:07amplified on hacker news 👑hn.story.49720533
sbulaev
peak 2 · 0 comments · 98% of case engagement
09-16 00:21our radar first saw it · +58.4hdiscovery anchor: hn.story.49720533—
pace: p23 vs 1032 stories at the 336h mark (now 674h old) — ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDivide, Consult, Conquer: Capability Laundering Through Aligned LLMs
Retrieved article excerpt

Open article · Retrieved 2026-09-16T00:22:20.941836+00:00

# Computer Science > Cryptography and Security

**arXiv:2609.15383** (cs)

[Submitted on 14 Sep 2026]

# Title:Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Authors:[Mark Russinovich](https://arxiv.org/search/cs?searchtype=author&query=Russinovich,+M), [Blake Bullwinkel](https://arxiv.org/search/cs?searchtype=author&query=Bullwinkel,+B), [Giorgio Severi](https://arxiv.org/search/cs?searchtype=author&query=Severi,+G), [Cristian Ovadiuc](https://arxiv.org/search/cs?searchtype=author&query=Ovadiuc,+C), [Ahmed Salem](https://arxiv.org/search/cs?searchtype=author&query=Salem,+A)

View a PDF of the paper titled Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs, by Mark Russinovich and 4 other authors

[View PDF](https://arxiv.org/pdf/2609.15383)
[HTML (experimental)](https://arxiv.org/html/2609.15383v1)
> Abstract:Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally. We call this attack capability laundering. Unlike a jailbreak, no single response is a harmful task. We measure consultation-aided uplift using tasks that a raw frontier model solves, the aligned frontier refuses, and the unassisted orchestrator fails. We evaluate GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants to four local orchestrators on CyBench, BountyBench, and harmful CBRN requests. On CyBench, Gemma-4-31B recovers 8/14 candidates with GPT-5.5 and 7/9 with Opus, compared with 2/21 and 4/15 for Gemma-4-12B. On BountyBench, Gemma-4-31B recovers 3/9 and 2/3 candidates, while Muse-Glimmer-30B recovers none of 22 and 13. For CBRN, we measure uplift across eight steps of a hypothetical bioweapon attack chain and find that consultation raises Gemma-4-31B's mean rubric score from 62.3 to 83.1 on a 100-point rubric scale. These results expose a gap in current defenses: refusing a harmful task does not prevent frontier capabilities from being transferred and composed across many individually permitted interactions.

|  |  |
| --- | --- |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | [arXiv:2609.15383](https://arxiv.org/abs/2609.15383) [cs.CR] |
|  | (or  [arXiv:2609.15383v1](https://arxiv.org/abs/2609.15383v1) [cs.CR] for this version) |
|  | <https://doi.org/10.48550/arXiv.2609.15383> Focus to learn more  arXiv-issued DOI via DataCite |

## Submission history

From: Ahmed Salem [[view email](https://arxiv.org/show-email/2955d358/2609.15383)]   
 **[v1]**
Mon, 14 Sep 2026 11:10:57 UTC (675 KB)

Full-text links:

## Access Paper:

View a PDF of the paper titled Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs, by Mark Russinovich and 4 other authors

- [View PDF](https://arxiv.org/pdf/2609.15383)
- [HTML (experimental)](https://arxiv.org/html/2609.15383v1)
- [TeX Source](https://arxiv.org/src/2609.15383)

[license icon](http://creativecommons.org/licenses/by/4.0/ "Rights to this article")

### Current browse context:

cs.CR

[< prev](https://arxiv.org/prevnext?id=2609.15383&function=prev&context=cs.CR "previous in cs.CR (accesskey p)")
  |   
[next >](https://arxiv.org/prevnext?id=2609.15383&function=next&context=cs.CR "next in cs.CR (accesskey n)")

[new](https://arxiv.org/list/cs.CR/new)
 | 
[recent](https://arxiv.org/list/cs.CR/recent)
 | [2026-09](https://arxiv.org/list/cs.CR/2026-09)

Change to browse by:

[cs](https://arxiv.org/abs/2609.15383?context=cs)  
[cs.AI](https://arxiv.org/abs/2609.15383?context=cs.AI)

### References & Citations

- [NASA ADS](https://ui.adsabs.harvard.edu/abs/arXiv:2609.15383)
- [Google Scholar](https://scholar.google.com/scholar_lookup?arxiv_id=2609.15383)
- [Semantic Scholar](https://api.semanticscholar.org/arXiv:2609.15383)

export BibTeX citation
Loading...

## BibTeX formatted citation

×

loading...

Data provided by:

### Bookmark

[BibSonomy](http://www.bibsonomy.org/BibtexHandler?requTask=upload&url=https://arxiv.org/abs/2609.15383&description=Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs "Bookmark on BibSonomy")
[Reddit](https://reddit.com/submit?url=https://arxiv.org/abs/2609.15383&title=Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs "Bookmark on Reddit")



Bibliographic Tools

# Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer *([What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))*

Connected Papers Toggle

Connected Papers *([What is Connected Papers?](https://www.connectedpapers.com/about))*

Litmaps Toggle

Litmaps *([What is Litmaps?](https://www.litmaps.co/))*

scite.ai Toggle

scite Smart Citations *([What are Smart Citations?](https://www.scite.ai/))*

Code, Data, Media

# Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv *([What is alphaXiv?](https://alphaxiv.org/))*

Links to Code Toggle

CatalyzeX Code Finder for Papers *([What is CatalyzeX?](https://www.catalyzex.com))*

DagsHub Toggle

DagsHub *([What is DagsHub?](https://dagshub.com/))*

GotitPub Toggle

Gotit.pub *([What is GotitPub?](http://gotit.pub/faq))*

Huggingface Toggle

Hugging Face *([What is Huggingface?](https://huggingface.co/huggingface))*

ScienceCast Toggle

ScienceCast *([What is ScienceCast?](https://sciencecast.org/welcome))*

Demos

# Demos

Replicate Toggle

Replicate *([What is Replicate?](https://replicate.com/docs/arxiv/about))*

Spaces Toggle

Hugging Face Spaces *([What is Spaces?](https://huggingface.co/docs/hub/spaces))*

Spaces Toggle

TXYZ.AI *([What is TXYZ.AI?](https://txyz.ai))*

Related Papers

# Recommenders and Search Tools

Link to Influence Flower

Influence Flower *([What are Influence Flowers?](https://influencemap.cmlab.dev/))*

Core recommender toggle

CORE Recommender *([What is CORE?](https://core.ac.uk/services/recommender))*

- Author
- Venue
- Institution
- Topic


About arXivLabs

# arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).

[Which authors of this paper are endorsers?](https://arxiv.org/auth/show-endorsers/2609.15383) |
Disable MathJax ([What is MathJax?](https://info.arxiv.org/help/mathjax.html))
sbulaev20
🟧 echo.paper ⭐The authors report consultation-aided uplift on cybersecurity benchmarks and harmful CBRN requests when local orchestrators decompose tasks Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, and Ahmed Salem——

Interpretation history

Decision trace