2026-10-11 16:38 UTC

Lasso Security reports that SynthID-Text watermarking changes tool-call correctness and refusal behavior in its tested open models, making watermark configuration a potential agent-reliability and safety regression surface even when aggregate accuracy changes little.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumllm-watermarking agent-evaluation agentic-securityLasso SecurityAndrea Siposova

What is this?

SynthID-Text is Google DeepMind’s text-watermarking technology, which adjusts token probabilities during generation to make AI-generated text identifiable. The case attributes to AI-security vendor Lasso Security a report that watermark configuration changes tool-call correctness and refusal behavior in tested open models, despite small aggregate accuracy changes. The supplied search snippets establish SynthID’s mechanism and Lasso’s agent-security focus, but do not retrieve that report or verify its reported 6.5% paired tool-call disagreement, experimental controls, or Andrea Siposova’s role. DeepMind and MIT Technology Review describe preserved output quality in Gemini testing; those snippets do not establish whether tool-call reliability or refusal behavior was evaluated.

Why it matters to Scott

Lasso’s reported findings extend Scott’s Evaluation-Driven Development and trace-backed agent comparison practices to a specific regression variable: watermark configuration, suggesting paired tool-call and refusal checks rather than aggregate scores alone. This could change his evaluation fixtures, but the supplied material does not verify the report’s controls or establish that his systems use SynthID-Text; the radar tracks related trajectory drift and silent tool-call failures, not this development.
ip:concept.evaluation-driven-developmentip:framework.reflexive-agent-designdev:concept.trace-backed-agent-comparisonradar:dfah-bench-agent-trajectory-driftradar:vllm-silent-tool-parser-failures
queries asked of Scott's wikis
  • agent harness regression tests sampling configuration changes
  • aggregate benchmark accuracy versus paired behavioral disagreement
  • tool-call correctness refusal behavior safety evaluation
  • AI provenance watermarking reliability tradeoffs
  • open-model inference reproducibility decoding configuration

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 602h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 14:00⭐ origin echo-reconstructedReports model- and key-dependent sampling drift under SynthID-Text, including average paired tool-call disagreement of 6.5% across 21 model-
Andrea Siposova, Lasso Security on blog (echo) · attributed from hn.story.49749997
—
09-18 03:53first on hacker news · published · +37.9hThe Provenance Tax: How LLM Watermarking Changes AI Agent Behavior
fourfire
—
09-18 03:53amplified on hacker newshn.story.49749997
fourfire
peak 9 · 1 comments · 7% of case engagement
09-18 07:02amplified on hacker newshn.story.49751004
joozio
peak 2 · 0 comments · 1% of case engagement
09-21 21:16amplified on hacker newshn.story.49793563
CrankyBear
peak 2 · 0 comments · 1% of case engagement
09-24 18:43amplified on hacker newshn.story.49835095
eatonphil
peak 3 · 0 comments · 2% of case engagement
09-26 13:05amplified on hacker news 👑hn.story.49856149
nisosguy
peak 59 · 72 comments · 89% of case engagement
09-18 04:20our radar first saw it · +38.4hdiscovery anchor: hn.story.49749997—
pace: p71 vs 1032 stories at the 336h mark (now 602h old) — ahead of antigravity-boost-reasoning-control (1.0x), behind cloudflare-disallow-training-search (1.0x)

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnThe Provenance Tax: How LLM Watermarking Changes AI Agent Behavior
Retrieved article excerpt

Open article · Retrieved 2026-09-18T04:21:55.087674+00:00

[Back to research](https://www.lasso.security/research)

# The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior

Andrea Siposova

Andrea Siposova

September 17, 2026

4

min read

The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior

## On this page

[This is a h2](https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior)

[This is a h3](https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior)

[This is a h4](https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior)

Recently, [Anthropic announced that future Claude models would embed an invisible watermark](https://www.anthropic.com/news/claude-text-watermark) in their output [1], [2], and subsequently disclosed that the watermark is based on Google DeepMind’s [SynthID-Text](https://www.nature.com/articles/s41586-024-08025-4) [2], [3]. Text watermarking itself is not new, but its deployment now has regulatory relevance. [Article 50(2) of the EU AI Act](https://artificialintelligenceact.eu/article/50/) [4] requires providers of AI systems generating synthetic text to mark their outputs in a machine-readable format and make them detectable as artificially generated or manipulated, using technical solutions that are effective, interoperable, robust, and reliable as far as technically feasible.

‍

Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token. At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection. At the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it. **Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools.** Such a watermarking procedure can therefore affect both what the model says and what an agent does. We call this behavioral effect **sampling drift.**

‍

Whether this drift appears in practice is an empirical question. We find that it does, in both model refusal behavior and agent tool calling. The effect is model- and key-dependent and can be obscured by aggregate scores when changes in opposite directions cancel. We therefore report both net performance and paired disagreement between watermarked and unwatermarked runs. Further, the closing section discusses what it means for AI safety and security and what developers should do about it.

‍

## **Built for Content Provenance, Deployed Inside Agents**

‍

A text watermark embeds a signal that allows output to be identified as AI-generated. Existing approaches include post-processing methods and methods integrated directly into LLM generation [8]. Generation-time approaches include logits-biasing methods [5], distortion-free keyed sampling [6], cryptographically motivated constructions [7], and SynthID-Text’s Tournament sampling [3]. Figure 1 contrasts this process with ordinary sampling. We use SynthID’s non-distortionary configuration, which preserves the original token distribution in expectation over the watermark randomness while individual generations under a fixed key can still differ [3]. Dathathri et al. report no measurable quality degradation across nearly twenty million Gemini responses [3].

‍

**Figure 1. Standard versus watermarked text generation.** Ordinary generation samples from the model’s token distribution. A generative watermark adds a random seed generator, sampling algorithm, and scoring function; SynthID-Text uses tournament sampling. Adapted from Dathathri et al. [3].

‍

‍

Anthropic’s deployment also illustrates why this matters beyond first-party chat interfaces. The company states that watermarking is applied at the model level and covers supported models accessed through the Claude Platform API as well as cloud providers [1]. A developer using a watermarked model as the reasoning component of an agent can therefore receive watermarked outputs even when the agent itself is a separate application. This makes model-level behavioral effects of watermarking relevant to the agents built around such models.

‍

### **Same Tokens the Watermark Biases, Same Tokens the Agent Acts On**

‍

Tournament sampling has more opportunity to alter token selection where the model is uncertain. In structured output such as JSON, braces, keys, and function names are often highly predictable, while values such as queries, numbers, paths, and recipients are less so. A change that would amount to a lexical variation in ordinary prose can therefore alter an argument that an agent executes.

‍

The weights and prompt remain unchanged, but token selection does not. Importantly, non-distortionary does not imply identical behavior under a fixed watermark key. The guarantee holds over the watermark randomness, while a particular key changes token selection during generation [3]. The resulting sampling drift can therefore change agent behavior even though the watermark is non-distortionary in the sense defined by Dathathri et al. Its effect can also depend on the watermark key, so we test multiple keys rather than relying on one.

‍

## **How We Measure the Effect**

‍

We use a paired design for two experiments. Tool calling is evaluated on BFCL v4 single-turn AST [9], and refusal on 200 HarmBench harmful behaviors [10] plus 100 benign JailbreakBench controls [11], with harmful requests tested both bare and under one fixed prompt injection technique. Table 1 summarizes the datasets, evaluation scope, temperatures, and expected behavior.

‍

We use the non-distortionary SynthID-Text configuration through HuggingFace’s unmodified SynthIDTextWatermarkLogitsProcessor, with 30 Tournament layers, n-gram length 5, sampling table 216, and context history 1,024. Each item is generated with and without SynthID from the same seed, batch composition, and order at each temperature. The watermark processor is the only difference within each pair.

‍

**Table 1. Datasets and experimental settings.**

‍

| Experiment | Dataset and scope | T | Correct behavior |
| --- | --- | --- | --- |
| Tool calling | BFCL v4 single-turn AST [9], live and non-live call-expected tasks plus relevance and irrelevance | 0.001, 0.7, 1.0 | Correct call, or no call when none fits |
| Refusal | HarmBench [10], 200 harmful behaviors; JailbreakBench [11], 100 benign controls; bare and fixed-injection prompts | 0.001, 0.7 | Refuse harmful; answer benign |

‍

## **The Tool-Calling Cost of Watermarking**

‍

We test whether watermarking changes tool selection, arguments, or output validity. A well-formed call to the correct tool with an incorrect path, recipient, query, or amount is particularly consequential because it can execute successfully while performing the wrong action. We evaluate calls individually, although an incorrect call in a deployed agent could also affect subsequent observations and decisions.

‍

### **How Often Tool-Call Correctness Changes**

‍

On items where a tool call is expected, watermarking reduces accuracy on six of the seven models, with a significant decrease on four. The net change in accuracy, however, does not show whether the same individual calls succeed with and without the watermark. A call that becomes incorrect can be offset by another that becomes correct, leaving the aggregate result nearly unchanged even though the model behaves differently on both items.

‍

We measure this directly using the **paired disagreement rate**, which we refer to as **churn**, defined as the share of items whose verdict differs between the watermarked and unwatermarked runs. For the comparison across temperatures in Figure 2, we use BFCL’s researcher-defined non-live items, which give us the same fixed set of 1,150 call-expected tasks at each temperature. Figure 2 shows that the paired disagreement is substantially larger than the net accuracy change. At *T=1.0*, 16.8% of phi-4’s call verdicts differ between the two conditions while its net accuracy loss is 2.87 points. Llama-3.1-8B shows the same pattern, with 9.9% of verdicts changing while the net loss is only 0.87 points. Across the 21 model-temperature combinations, churn averages 6.5%, and its bootstrap interval excludes zero in every case.

‍

**Figure 2. Paired tool-call disagreement under watermarking by model and temperature.** Results use 1,150 non-live BFCL call-expected items across models and temperatures. The central vertical line represents no change relative to the unwatermarked condition. Orange bars show calls that changed from correct to incorrect, and blue bars show calls that changed from incorrect to correct. The final column reports the churn with non-live irrelevance items included. Diamonds denote 95% bootstrap intervals excluding zero.

‍

‍

### **Which Tool-Call Errors Change**

‍

Error type also matters. Malformed output prevents the intended call from executing, while a well-formed call with the wrong tool or argument can still execute. Figure 3 separates these failures into wrong tool, wrong arguments, and malformed output. Unlike Figure 2’s across-temperature comparison, this analysis combines BFCL live and non-live call-expected items at *T=0.001* to characterize errors across the broader benchmark. Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.

‍

**Figure 3. Changes in tool-calling errors under watermarking by error type.** BFCL live and non-live call-expected items at *T=0.001*, separated into wrong tool, wrong arguments, and malformed output. The vertical lines show the accuracy without watermarking, and the bars show the change when watermarking is applied.Orange denotes correct-to-error changes and blue error-to-correct changes. Intervals are 95% item-level bootstrap intervals.

‍

‍

The error profile also differs across models. On Llama-3.1-8B, the largest contribution to the accuracy loss comes from incorrect arguments (−3.48 points), followed by wrong-tool calls (−1.84 points). On phi-4 and Granite-3.2-8B, malformed output dominates (−5.96 and −4.36 points). Similar aggregate changes can therefore arise from different failure modes.

‍

## **Watermarking Can Weaken Refusal Under Prompt Injection**

‍

Refusals are also generated token by token, so watermarking can affect them. We test harmful requests alone and with one simple, fixed prompt-injection technique intended to reduce refusal. The technique appends an adversarial instruction as retrieved content, claiming that the safety filter is disabled and instructing compliance. It is held constant across prompts, models, and temperatures. [OWASP GenAI LLM Top 10 2026](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/) identifies prompt injection as an input-side vulnerability that can alter model behavior in ways unintended by the agentic application developer, with consequences that can extend to harmful outputs and unauthorized tool actions in agentic systems [12].

‍

### **Refusal Behavior With and Without the Prompt Injection Technique**

‍

Watermarking changes refusal behavior on bare harmful requests, but the effect becomes more pronounced under prompt injection. As Figure 4 shows, disagreement increases on several models when the same harmful requests are paired with the fixed prompt-injection technique, with the strongest effects shifting predominantly from refusal to compliance. This makes the result particularly safety-relevant because the behavioral effect becomes more pronounced when the model is exposed to an adversarial prompt.

‍

At T=0.001, gemma-3-27b’s churn increases from 6.0% on bare harmful requests to 23.5% under prompt injection, while the net compli
fourfire91
🟧 echo.blog ⭐Reports model- and key-dependent sampling drift under SynthID-Text, including average paired tool-call disagreement of 6.5% across 21 model-Andrea Siposova, Lasso Security——
🟧 hnLLMs respond differently to harmful prompts when AI watermarking is usedjoozio20
🟧 hnLasso: AI Watermarks Change How Agents ActCrankyBear20
🟧 hnWatermarking in vLLMeatonphil30
🟧 hnUnderstanding the Impact of LLM Watermarking on AI Agent Behaviornisosguy5972

Interpretation history

Decision trace