Retrieved article excerpt
Open article · Retrieved 2026-09-16T17:24:50.271780+00:00
[THE SURGE AI BLOG](https://surgehq.ai/blog)
## We Trained a Model on Office Work. It Also Got Better at Coding.
September 16, 2026
August 3, 2026
Surge AI Research Team
TL;DR
We trained a model on office tasks, none involving code. It still gained **+5.8pp on SWE-Bench Pro**, which raises a question: how does learning to fill out a spreadsheet help a model fix a PowerShell parser?
Analysis of the model trajectories points to a general capability we call **Goal-Directed Execution**: the ability to form appropriate goals, build an accurate understanding of an unfamiliar environment, maintain higher-level objectives through complex work, and verify that the intended outcome was achieved.
High-quality tasks do more than teach domain-specific knowledge. They strengthen higher-level capabilities that transfer to entirely new domains.
[Research Paper](https://arxiv.org/abs/2608.01604)
Table of contents
[Case Study Llama](https://surgehq.ai/blog/office-work-post-training-improves-coding)[Analysis of Individual writings](https://surgehq.ai/blog/office-work-post-training-improves-coding)
[Appendix](https://surgehq.ai/blog/office-work-post-training-improves-coding#appendix)
We post-trained Qwen3.5-122B-A10B on a suite of RL environments for long-horizon work involving documents, spreadsheets, web research, planning, and tool use across office services. **None of the training tasks involved coding. The model got better at coding anyway.**
It improved on held-out tasks from the same distribution, which is what you'd hope for:
Corrected heldout table
It [generalized](https://surgehq.ai/blog/cross-benchmark-generalization-for-long-horizon-agentic-tasks) to other tool-use benchmarks the training data never targeted:
Corrected tool-use table
It also improved on software-engineering benchmarks it had never seen:
\_\_wf\_reserved\_inherit
The post-training data supplied no repository conventions, parser semantics, code-level solutions, software-specific graders, or benchmark feedback, and SWE-Bench Pro still improved by 5.8 percentage points. Which raises the obvious question: **what does filling out a spreadsheet teach a model that helps it fix a PowerShell CLIXML parser?**
Our answer: **the training taught the model to form the right goals, keep them stable under pressure, and continuously check its picture of the environment against them.** Those skills transfer because every agent runs the same underlying loop, whether it's updating a workbook or modifying a codebase.
Picture a junior engineer who takes a year off to plan weddings. When they return, their code reviews get noticeably better. Nothing about seating charts teaches them Python. What the year teaches them is how to run a complex project where everything depends on everything else:
- Turn "make it magical" into a concrete plan with dates, vendors, and a budget (translate a loose ticket into a precise change to make)
- Track how one change ripples through the rest: a new venue breaks the caterer's timeline (a change to one module ripples into every consumer that depends on it)
- Keep the couple's actual wishes in view while firefighting on the day (don't let a stubborn bug pull you into changes that break the actual requirement)
- Do the final walkthrough before the guests arrive, not after (run the tests, then check what else depends on what you touched)
Those habits are exactly what you need to work in an unfamiliar codebase, and they were learned without writing a line of code.
*A quick caveat before we dig in: this is a behavioral interpretation, not a claim about what is represented inside the model's weights. The explanation comes from aggregate trajectory statistics and manual review of dozens of paired trajectories.*
## What transferred: the model learned goal-directed execution
To make "forming and pursuing goals" concrete, we can describe any agent task using three elements:
- **Goal:** the desired state of the environment.
- **Action:** an interaction intended to move the environment toward that state.
- **Working state:** the agent's current understanding of the environment and its relationship to the goal.
An effective agent repeatedly runs the following loop:
1. Define the goal.
2. Select an action that should move toward it.
3. Observe the result.
4. Update its working state.
5. Compare the new state with the goal.
6. Stop if the goal has been achieved; otherwise, act again.
\_\_wf\_reserved\_inherit
We call this cycle the **goal loop**. People run the same loop when parallel parking: check the mirrors, turn the wheel, inch backward, check again, adjust, and stop when the car sits where it should. The hard part for an agent comes from scale: a real task involves hundreds of these loops, nested inside one another, and each has to stay connected to the purpose it serves.
Because an agent rarely begins with complete knowledge, each action serves two purposes: it may change the environment, and it also reveals information about the environment. Opening a file, running a test, querying a database, and inspecting a spreadsheet all double as experiments that refine the agent's understanding of what is true.
When we compared trajectories before and after training, the clearest difference was the reliability with which the trained model executed this loop across long, complicated tasks. New domain knowledge played little role.
## Complex tasks require hierarchical goal-directed execution
In practice, agents don't execute a single goal loop. They execute a hierarchy of them.
Consider a task like: *Review our software subscriptions and recommend which vendor contracts to renew, renegotiate, or cancel before the quarter ends.*
That instruction is too abstract to execute directly. Before the agent can act, it has to break the task down. A plausible decomposition:
1. Locate the contract folder and the finance workbook that tracks spend.
2. Determine the decision criteria: the budget target, usage thresholds, notice periods for cancellation.
3. Extract the active subscriptions and their renewal dates.
4. Pull usage data for each tool from its admin dashboard.
5. Normalize the figures so tools are comparable (monthly vs. annual billing, per-seat vs. flat pricing).
6. Match spend against usage and flag tools that are underused or redundant.
7. Investigate the edge cases: a tool with few logins but a critical integration, a contract with an auto-renewal clause.
8. Produce the renew / renegotiate / cancel recommendation.
Each step is a subgoal that runs its own copy of the loop, and many decompose further. "Pull usage data" splits into finding each tool's dashboard, exporting activity reports, handling the tools that offer no export, and merging everything into one comparable sheet, each of which bottoms out in individual tool calls. A modest-sounding office task expands into a tree several levels deep and dozens of loops wide.
\_\_wf\_reserved\_inherit
To see how complex this becomes in practice, explore the full decomposition of a successful trajectory from the holdout set here.
Explore the full task decomposition
Decomposition, however, is only half the problem. Every branch reads from or modifies the same underlying environment, so what one branch discovers or changes affects what every other branch should do. The agent has to fold low-level observations into one coherent working state and understand what each observation means for the immediate subgoal *and* for the objectives above it. That dual bookkeeping is where things go wrong:
- **A locally sensible action can be globally wrong.** Converting every contract to a monthly cost makes the comparison step easy, and hides the fact that one annual contract auto-renews next week, which is exactly the deadline the task is about.
- **A relevant fact can get lost between levels.** While pulling usage data, the agent notes that most activity runs through service accounts, then carries only the five human logins into the comparison, and a heavily used tool looks dormant.
- **A subtask can succeed while undermining a parent requirement.** Cancelling a cheap, barely-used tool closes the budget gap, and silently breaks the integration that a flagship tool the company is keeping depends on.
- **An intermediate result can look plausible while being incomplete.** A spend list built from the finance workbook is a perfectly coherent list. Nothing about the list itself reveals the subscriptions expensed on individual credit cards that never reached the workbook.
Our trajectory analysis surfaced four recurring points at which this process breaks down, each corresponding to a distinct capability:
\_\_wf\_reserved\_inherit
Failures in an office workflow and a software repository look very different on the surface. Underneath, we kept finding the same four behavioral issues in both.
\_\_wf\_reserved\_inherit
## Why this training data exercises general capabilities
A model can improve on a benchmark for narrow reasons: it memorizes a tool's calling convention, or learns a dataset quirk. That kind of learning doesn't travel. These tasks were designed to capture the complexity of real work, where the model has to manage a hierarchy of goals while piecing together an accurate picture of a complicated environment. Task creators targeted four properties:
- **Deep decomposition:** high-level objectives had to be broken into branching and sequential subtasks.
- **Parallel investigation and synthesis:** evidence needed for a decision was distributed across sources and tools.
- **Entangled constraints:** requirements interacted and couldn't be handled independently.
- **Long dependent chains:** later work depended on earlier actions and conclusions being correct.
The tasks spanned 27 non-software categories and many different tools. A tool-specific shortcut therefore helped on only a small slice of the data, while better goal-loop discipline paid off across the entire collection.
\_\_wf\_reserved\_inherit
## What it looks like in the trajectories
The examples below show the same underlying problem appearing in both a software-engineering task and a general tool-use task. In every example, the model failed before training and passed after.
\_\_wf\_reserved\_inherit
### 1. Defining the right goal
Long tasks break into smaller goals, often across several levels. At each point, the model has to choose what to do next by connecting its immediate context to the larger objective.
**Software engineering: SWE-Bench Pro, NodeBB selected fields.** The task was to let database callers request only selected fields. Both models found existing helpers that already implemented and tested the required field-selection behavior. The correct local goal was to route the new option through those helpers. The base model instead reimplemented the behavior across the backends and changed existing missing-object semantics. The trained model reused the established path.
**General tool-use: Toolathlon, A/B conversion rates.** The task was to evaluate how often homepage clicks led to store views. Both models saw the same column names and values, including the fact that clicks exceeded store views. The base model computed clicks / store\_views, produced conversion rates above 100%, and kept going with the formula anyway. The trained model interpreted "conversion" in the context of the task and used store\_views / clicks.
\_\_wf\_reserved\_inherit
### 2. Building goal-relevant state
Each low-level action reveals another piece of the environment. The model needs to combine those pieces into a coherent picture and understand what they mean for the larger task, not just the current subtask.
**Software engineering: SWE-Bench Pro, package-name validation.** The task was to make package-name validation reject reserved language keywords. The base model had just read the existing imports and validation structure, yet its new helper called a module that was never imported, faili