Asana Cuts Browser-Agent Costs 76x with GPT-6.1 Sol: A Workflow Study Driven by Codex

Published · AI Daily — AI-assisted deep research, methodology & disclosure

In a 144-run Asana study, an optimized GPT-6.1 Sol browser agent averaged $0.47 estimated cost and four minutes per run, 76x cheaper and 5x faster than the old Model B setup. Codex found it did not cache its growing page and screenshot history.

On October 9, 2026, OpenAI published a customer story about Asana. The headline numbers are hard to ignore. In a set of browser-agent tests, Asana cut estimated model cost by a factor of 76 and made runs five times faster. The optimized workflow ran on GPT-6.1 Sol. It averaged $0.47 in estimated model cost and about four minutes per run. The baseline was the original production setup on Model B. A simple calculation from those two figures puts the old cost at roughly $36 per run. One caution comes first. This is a vendor-published case. The costs are estimates. The comparison comes from a 144-run study that Asana designed. It is a serious engineering signal, but it is not yet an industry benchmark that any team can copy. The background is StackAI, a platform Asana acquired. Customers use it to build workflows without writing code. Those workflows send an agent to navigate websites, fill out forms and gather information. One run looks harmless. At Asana's scale, every small inefficiency is multiplied by a very large call volume. So Frank Hidalgo, PhD, the CTO of StackAI, set out to make the browser agent faster and cheaper. He did not lead a team through a manual audit. He directed GPT-6 Astra in Codex to investigate the agent, test improvements and compare the results. By his estimate, this work would have taken one to two months by hand. It took about a week. The diagnosis is the most instructive part of the story. Hidalgo first asked GPT-6 Astra to map the codebase and explain how the agent built each model request. It found that the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots. Consider what that means. A browser agent takes a step, then hands the model everything it has seen so far, plus a new screenshot. The longer the history, the more input the model must process at each step. If that part of the request cannot hit the cache, the same content is billed at full price again and again within one run. It also slows the time to first token each time. The waste builds up with every step, and total input grows faster than the step count itself. This is our reading of the mechanism, and it is a reasonable one. OpenAI's public excerpt does not list every change that followed, and we should not invent the missing items.

Now consider the study design. Asana ran 144 runs. They covered GPT-6.1 Sol and three other frontier models, called Model A, B and C in the text. The strength of this design is that model choice and workflow redesign sit in the same comparison. It also creates an interpretation problem that must be stated plainly. The 76x figure compares the optimized Sol workflow with the original production setup on Model B. Two variables moved at once. Crediting all 76x to the model would overstate what the model did. Crediting all of it to the workflow would ignore real differences in model price and speed. The excerpt shows no control that separates the two, and that is where a careful reader should push. The costs are also estimates, so they will shift when vendors change prices. And 144 runs is a modest sample, so the stability of the result needs more repetition to confirm.

The second layer of meaning is about how the work was done. Arnab Bose, Asana's CPO, described it as what teams of humans and agents look like in practice: an engineer set the direction, GPT-6 Astra ran the experiments, and the results went through Command to production. There is a clear line of labor in that sentence. Direction, trade-offs and the decision to ship stayed with people. The slowest parts, which are reading code, forming hypotheses, running comparisons and tidying data, went to the agent. When the marginal cost of an experiment falls, optimization no longer has to wait for a free week. It can become a routine habit. But the center of value then moves to evaluation. Does the team have a set of repeatable tasks? Does it have comparable metrics? Does someone read the results with care? Without those, faster experiments only produce more noise.

For the industry, the case sends three signals. First, competition among browser agents is shifting from whether the task can be done to what each run costs in dollars and minutes. A drop from tens of dollars to under one dollar is what makes thousands of enterprise runs per day plausible. Second, the cache hit structure should become a standing item in agent architecture reviews, above all for context that grows with every step. Third, one stated aim of the work was to let Asana offer customers more capable models. Lower cost buys room for capability. For readers, the sound move is to borrow the method rather than the number. Audit which part of your own requests is not cached. Run a controlled comparison on one fixed task set. Test the model change and the workflow change separately. Then report cost next to success rate. A 76x saving is impressive, but it only means something if the agent still does the job right.

Sources

FAQ

Does the 76x cost drop come from switching models?

Not on its own. The baseline is the original production setup on Model B. The comparison is the optimized workflow on GPT-6.1 Sol. Two changes are mixed: the model and the workflow. The public excerpt does not split their contributions, so readers should ask for that breakdown.

What exactly was wrong with the agent's caching?

According to OpenAI's account, GPT-6 Astra found that the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots. A browser agent resends that history at every step, so uncached input is billed again and again. The effect grows with the number of steps.

Can other teams expect the same result?

Treat it with care. This is a vendor-published case, the costs are estimates, and the sample is 144 runs that Asana designed itself. Task type, site structure and model pricing all matter. The safer lesson is the method: find which part of each request is not cached, then run a controlled comparison on the same task set.