Basis completes tax workbook 2x faster with GPT-6 Astra

Published · AI Daily — AI-assisted deep research, methodology & disclosure

GPT-6 Astra completed a 50-tab tax workbook twice as fast as GPT-5.6 Sol, and its stronger understanding of user intent gives Basis more confidence in real-world use.

Background and Context

In September 2026, OpenAI disclosed the first public vertical-industry benchmark for its latest reasoning model, GPT-6 Astra, through a real-world test conducted by tax software platform Basis. The evaluation centered on a complex, 50-tab tax workbook—a document type that typically bundles income statements, deduction schedules, depreciation tables, and credit calculations across dozens of interdependent sheets. Basis reported that GPT-6 Astra completed the entire workbook twice as fast as its predecessor, GPT-5.6 Sol, while simultaneously demonstrating a markedly stronger grasp of user intent. The task was not a simple text-generation exercise; it required cross-table data extraction, alignment with current tax regulations, multi-step arithmetic derivation, and orchestration of a sequential workflow that mirrors the cognitive load of a seasoned tax preparer.

Basis’s engineering team emphasized that the speed gain was accompanied by a significant reduction in manual corrections and repeated prompting. In prior deployments, even advanced models would frequently misinterpret the implicit logic linking disparate tabs, forcing human accountants to intervene and re-prompt the system. With Astra, the end-to-end task completion rate rose sharply because the model could infer the accountant’s overarching goal—such as “prepare the federal and state filings for this entity”—and autonomously determine which forms to populate, which cross-references to resolve, and how to allocate deductions across jurisdictions. The September announcement positions GPT-6 Astra as one of the earliest proofs that frontier reasoning models can move beyond conversational fluency into reliable, multi-step professional execution.

Deep Analysis

The twofold speed improvement stems from a fundamental re-architecture of the model’s reasoning pipeline. Compared with GPT-5.6 Sol, Astra incorporates a more efficient long-context memory compression mechanism and an adaptive reasoning-path planner. When confronted with a 50-tab workbook, the model no longer treats each sheet as an isolated prompt requiring external tool calls or piecemeal chain-of-thought fragments. Instead, it constructs the entire implicit tax-logic chain in a single, coherent pass—identifying dependencies among tabs, sequencing calculations according to Internal Revenue Code rules, and flagging anomalous entries that could indicate data-entry errors or compliance risks. This holistic approach eliminates the costly back-and-forth that plagued earlier versions, where attention decay over long contexts would cause the model to lose track of earlier tabs’ outputs.

Crucially, Astra elevates “understanding user intent” from shallow instruction-following to goal-driven autonomous planning. For example, when a user submits a workbook with the instruction “complete this year’s federal and state returns,” the model must independently decide which schedules are mandatory, which figures must be pulled from supporting tabs, and how to apportion deductions between federal and state calculations. Astra’s enhanced chain-of-thought reasoning achieves this by maintaining a dynamic task graph that updates as new information is processed, allowing it to recover gracefully if a required field is missing and to prompt the user only when truly necessary. The result is not merely faster output but a higher first-pass accuracy that directly translates into the observed doubling of overall workbook completion speed.

Industry Impact

For Basis, which serves small and medium-sized accounting firms, the performance leap carries immediate commercial weight. A complex tax workbook that previously demanded hours of human review—even with AI assistance—can now be processed in half the time with fewer interventions. This allows Basis to increase client throughput without a proportional rise in staffing costs, and it opens a path to evolve the platform from an assistive tool into a semi-automated filing service. Improved intent understanding also reduces the cognitive burden on accountants, who can redirect their expertise toward higher-value advisory work rather than data verification.

The demonstration intensifies competitive pressure across the AI and tax-software landscapes. OpenAI’s ability to deliver a measurable, task-specific speed advantage in a regulated professional domain challenges rivals such as Anthropic and Google DeepMind. Anthropic’s Claude series has long been prized for long-document comprehension and safety, but Astra’s concrete benchmark in tax workflow completion may sway enterprise buyers toward OpenAI’s API ecosystem. Meanwhile, incumbent tax-software giants like Intuit (TurboTax) and Xero, which are weaving AI into their products, face a build-or-partner dilemma: in-house models struggle to match the iteration velocity of frontier labs, likely accelerating partnerships or acquisitions of third-party reasoning models. For the accounting profession, the short-term effect is a welcome reduction in repetitive manual labor, but the medium-term outlook suggests a role transformation—from data entry and cross-checking to strategic tax planning and client consultation.

Outlook

Several signals will determine how broadly Astra’s tax-workbook success translates to other industries. OpenAI is expected to release additional benchmarks for similarly structured, long-document reasoning tasks—such as audit workpapers, legal contract review, and insurance claims adjudication—to establish Astra’s generality. For Basis, the critical next step is to demonstrate that the speed and accuracy gains observed in testing persist in live commercial deployments, with published data on error rates and the frequency of human overrides. Transparent metrics will be essential to convert early adopter enthusiasm into sustained enterprise trust.

Competitor responses are already taking shape. Anthropic may accelerate a professional-document-optimized version of Claude, while Google could deepen the integration of Gemini’s reasoning capabilities with Google Sheets, creating a native productivity loop. More broadly, GPT-6 Astra’s performance marks a turning point in the evolution of large models: the value proposition is shifting from generating fluent text to reliably executing multi-step, rule-intensive professional workflows. This transition will accelerate AI adoption in finance, law, and healthcare, but it also raises the stakes for model explainability, auditability, and regulatory compliance. The Basis case is an early validation that reasoning models can substitute for segments of human professional labor, and the pace of that substitution may outstrip current industry expectations.

Sources