What to Ask an AI Tax Vendor Before Tax Season

Short answer: Evaluate any AI tax vendor on four things: its tax engine, its rollout model, how it catches and traces errors, and the commercial path.
When to evaluate AI tax software before tax season
The Q4 evaluation window after the September 15 and October 15th extension deadlines is when most firms make platform decisions for TY26. Firms that start a controlled pilot now will have evidence before January. Firms that wait until January start from zero.
What makes AI tax preparation vendor evaluation harder than other software purchases is the level of stakes involved. A tax preparation platform sits at the center of client relationships, reviewer workflows, and signer liability. A demo shows the interface working. It does not show whether the platform holds up when a senior reviewer opens a 1065 on a September deadline with a filing clock running.
Knowing what to ask AI tax preparation software vendor teams before a demo is what separates a useful evaluation from a guided tour. The ten questions below are organized by what they test: the technical foundation, the implementation model, the risk controls, and the commercial structure. Most vendors will have practiced answers for all of them. The follow-up questions are where the evaluation gets useful.
What an AI tax preparation vendor must prove
The most important question to ask is whether the platform is approved for IRS e-file and MeF compliant, calculating against a governed compliance system, or whether it only generates output that still has to be checked against one.
Does this platform have a dedicated Tax Engine, or is it AI layered on top of legacy software?
The most important technical question in an AI tax vendor evaluation is whether the platform has a Tax Engine, a purpose-built compliance system that governs calculations, forms logic, and filing paths, or whether its AI is a layer on top of software not designed for AI-native operation.
An AI agent without a Tax Engine in the loop is generating outputs rather than calculating them. It does not know what a specific K-1 line requires. It cannot enforce the rules that determine whether a deduction applies. It cannot execute an e-file path with government authorization. What it can do is produce text that looks like a return, which creates a different kind of risk than a traditional software error because the failure mode is less visible.
Instead's AI agents work against a Tax Engine, the system that governs calculations, compliance logic, and filing paths. The research the platform surfaces is grounded in the IRC, Treasury Regulations, and court cases, not general training data. The buyer's question is simple: does the vendor's AI calculate against a real compliance engine, or does it generate answers that then need to be checked against one?
Which entity types and jurisdictions is the vendor actually approved to e-file today?
Ask the vendor to state, explicitly, which entity types and jurisdictions it is approved to e-file, and to separate two things that are often conflated: approval and calculation coverage versus live in-product filing.
A vendor may be able to populate a return for a given jurisdiction without being authorized to transmit it electronically. A vendor may hold IRS approval to e-file Individual returns but not Partnership or Trust returns. The IRS Authorized e-file Providers program is public; government authorization to e-file is a regulatory status, not a marketing claim. Ask the vendor to show you the specific entity types and jurisdictions where it has been approved and where it is actively filing, and to tell you clearly which are which.
Instead holds government authorization for e-filing across federal and state jurisdictions for five entity types: Individuals (1040), C Corporations (1120), S Corporations (1120-S), Partnerships (1065), and Trusts and Estates (1041). Approval must be verifiable by form and jurisdiction; it should never be presented as government endorsement. For a worked example of one entity type end-to-end, see how 1065 filing works with Instead. For any vendor, the right question is not "how many jurisdictions," but rather: show me the current approval scope, and tell me which of those are live for clients today. Calculation capability and authorization to transmit are different things, and every reputable vendor should be able to distinguish them.
Is the AI accessing live IRS publications or working from training data?
Tax law changes. An AI system trained on IRC sections and Treasury Regulations as of a fixed date can produce plausible outputs based on superseded guidance. The evaluation question is not whether the AI knows tax law; it is whether the platform's research outputs are grounded in the current regulatory record.
The right answer describes a citation structure: where the AI sources its regulatory references, whether those sources are live authoritative collections, and whether the outputs surface the specific provision that each answer draws from. A system that cites "general tax knowledge" without a traceable provision is a different product from one that shows the exact IRC section or Revenue Procedure behind each answer. For a tax professional who will sign the return, that distinction is not optional.
What rollout of an AI tax platform actually looks like for your firm
The implementation questions test whether the vendor is offering a structured control path or asking for a firmwide commitment before the platform has proved itself.
Is this a day-one replacement or a pilot that earns trust before full replacement?
Most vendor discussions frame the decision as: do you switch or not? The better framing is: What do the first 90 days look like operationally, and what evidence does the firm have at the end of those 90 days?
The strongest implementation models for large firms do not begin with a firmwide cutover. They begin with a controlled lane, one return type, a defined cohort of source documents, a reviewer gate, and a signer gate, that runs alongside the legacy system until the new platform has proved itself on representative work. Leadership gets evidence before making an irreversible operating decision.
For a Top 200 firm, implementation is part of the control system. It defines the first cohort, the workflows that need customization, the evidence reviewers need, the signer gate, and when the firm can expand or stop. A vendor presenting implementation as a services package rather than a structured control path is describing a different kind of risk. The polished demo is not the risky part. The cutover year is.
What changes for your team's workflow on day one?
The first-day experience determines whether the pilot generates usable evidence or just friction. Ask the vendor to walk you through what a preparer does differently on day one, not a demo, but a walkthrough of the actual workflow: where source documents go, how the workpaper is produced, where the reviewer enters the return, and what they find when they open it. For a step-by-step example of that path on an Individual return, see the 1040 filing workflow in Instead, intake to e-file.
The reviewer's experience is worth pressing on specifically. Ask what the reviewer opens: a workpaper with source citations for each value and findings already organized by severity, or a return that must be assembled before it can be assessed. In a legacy workflow, the reconstruction that precedes judgment is billed at reviewer rates. That is where the difference shows up.
How an AI tax platform finds, corrects, and traces errors
The risk questions test whether errors are findable, correctable, and distinguishable from acceptable judgment calls before a return is filed.
When the system makes a mistake, how is it surfaced and corrected before filing?
Every AI system makes mistakes. The evaluation question is not whether errors occur; it is whether the platform's design makes errors findable, correctable, and distinguishable from acceptable judgment calls before the return is filed.
A structured review workflow does this by organizing findings before the reviewer opens the file. In Instead's two-tab AI Review structure, every check the review pass runs is organized by severity, Critical, High, Medium, and Low, and every finding surfaces the specific figure involved, the section it came from, and its status. The reviewer's judgment sorts each finding into one of three categories: a real error requiring correction at the source, an acceptable position they confirm, or a false positive they dismiss. Separately, the Addressed field records Yes, No, or N/A on whether the finding was addressed.
Ask any vendor how its review workflow distinguishes between those three dispositions and how the distinction is recorded. A workflow that flags items but does not capture the reviewer's disposition creates a different audit record than one that requires a documented decision on every flagged item. For a deeper look at how the two-tab review structure works in practice on a 1065 return, see how AI tax return review works in Instead.
Is there an audit trail that traces every output back to its source document?
An audit trail in tax preparation means something specific: the ability to trace a value on the return back to the source document that produced it, without manually reconstructing the path. Schedules built from many underlying records, such as fixed assets and the depreciation that feeds the return, are where that requirement is tested hardest.
In Instead's workpaper, every value carries a recorded reference to its source document and workpaper location, so a reviewer can see where a number came from without rebuilding the mapping. When a partner asks why a specific item was not flagged, the answer is in the AI Review - Passed tab. When a reviewer hands off mid-engagement, the incoming reviewer can see what was already dispositioned and what remains.
The question in any vendor evaluation is whether the workpaper structure makes this connection explicit at each value or leaves the reviewer to establish it during review. The answer determines how much of a reviewer's session is judgment and how much is reconstruction.
What the AI tax software buying decision involves
The commercial questions test whether the pricing, pilot structure, and accountability model match how a large firm actually evaluates and adopts a platform.
How is pricing structured: per return, per seat, or subscription?
The pricing structure determines whether the platform's economics align with your volume and your workflow. A per-return model creates a different incentive structure than a per-seat subscription, particularly for firms with variable seasonal volume. Ask how pricing scales across entity types, whether complex returns such as a 1065 with K-1s carry the same cost structure as simpler ones, and what pricing looks like during a parallel license period.
The parallel license period is a real cost of evaluation. A firm running an AI tax platform alongside its existing software carries two license costs during the pilot window. A vendor that does not address this in pricing discussions is not accounting for how large firms actually evaluate new platforms.
What does a proof-of-concept path look like before firmwide rollout?
The right proof-of-concept path starts narrow: one return type, a cohort of source documents the review team knows well, defined evidence criteria before the pilot begins, and a clear decision point for expansion. A vendor that cannot describe this path in operational terms, what return types, what approval criteria, what evidence the firm will have at 30 and 60 days, is asking for a firmwide commitment without the structure to de-risk it.
The firms that evaluate AI tax platforms successfully go into the pilot with the evidence criteria already defined. What does the reviewer need to see to recommend expansion? What would cause the firm to stop at 60 days? For a detailed look at how a governed rollout works operationally, see how Top 200 firms should plan a safe AI-native tax platform rollout.
What happens if the system makes a mistake on a filed return?
This question tests whether the vendor has thought through the accountability structure, not just the error-prevention design. An AI-prepared return that contains an error after e-filing creates a real workflow event: an IRS notice, an amended return, and potential penalty exposure.
Ask which party bears responsibility for a filed error and what the remediation path looks like. Ask whether the platform surfaces the relevant authorization forms, Forms 8879, 8879-PE, and 8879-S, before transmission. The IRS e-file authorization structure imposes specific obligations on the Electronic Return Originator under IRS Publication 1345 for Individual returns and IRS Publication 4163 for Business returns, and the vendor should explain where those obligations sit in their workflow. For context on what filing authorization means in practice, see what filing approval means for an AI-native tax system.
Five questions to ask before committing to an AI tax vendor
The ten questions above produce a lot of information. These five filter the most important decisions before the firm commits team capacity:
- Can the vendor provide a current-state list of the entity types and jurisdictions in which it is approved to e-file, separated from jurisdictions where it calculates but cannot transmit?
- Does the vendor have a Tax Engine that the AI works against, or is the AI generating outputs without a governed compliance path?
- Can a reviewer trace any value back to the source document that produced it, without reconstructing the path?
- Does the implementation model describe a governed lane alongside the existing system, or require a firmwide commitment first?
- Is the pricing structure transparent about the parallel license period, and what evidence criteria support the expansion decision?
How to see Instead before committing team capacity
The Instead AI tax platform covers tax preparation, review, and filing across Individual and Business returns. The AI review stage sits within an integrated workflow that runs from source document intake through e-file authorization in a single path, rather than across separate tools connected by manual exports.
The ten questions above are not Instead-specific. They are how to evaluate AI tax software before tax season, whichever vendor a firm is considering, and they matter most before Q4 closes and January planning begins. Walk through the Instead platform with our team to see how Instead answers them against your own return mix.
Request a vendor evaluation session.
Frequently asked questions
Q: What is the difference between an AI tax agent and AI tax software?
A: AI tax software is a broad term for any platform using artificial intelligence in tax preparation or filing. An AI tax agent is more specific: a system designed to act on instructions, work through multi-step preparation tasks, and adapt to the source documents it encounters. What matters in a firm evaluation is whether the agent works against a real tax engine that governs calculations, forms logic, and e-file paths. An agent with one produces outputs that run through a compliance path. An agent without one produces outputs that only look like they ran through a compliance path.
Q: What is a Tax Engine, and why does it matter for tax preparation?
A: A Tax Engine is the part of a tax platform that handles calculations, enforces forms logic, and governs the filing path. When an AI agent works against one, its outputs are checked against actual compliance rules at each step. When AI is layered on legacy software not built for AI-native operation, the agent produces outputs that the legacy engine may not process correctly, and the reviewer catches the gap. The difference shows up in what the reviewer opens: a workpaper they can verify, or output they still have to reconstruct.
Q: How long does a proof-of-concept pilot typically take before a firm can make an expansion decision?
A: A well-structured pilot can produce meaningful evidence within 30 to 60 days. The evidence criteria matter more than the timeline: can the reviewer verify that every value traces to its source, do the findings surface what the firm's senior reviewer would surface, and does the e-file path match the firm's authorization workflow? A vendor that cannot describe those criteria before the pilot begins is offering a trial period rather than a governed evaluation.
Q: What should a firm look for in a vendor's error-handling model?
A: Three things: how errors are surfaced before filing, how the reviewer records what they did with each finding, and where corrections flow. A system that organizes findings by severity before the reviewer opens the file is a different model from one that relies on the reviewer to find errors in the output. A system that records both the reviewer's judgment on each item and whether it was addressed keeps a different engagement record than one that flags items without a documented decision. Routing corrections back to the workpaper source is a different risk profile from allowing direct edits.
Q: What does "government-authorized e-filing" mean for a firm evaluating an AI tax platform?
A: IRS e-file authorization means the IRS has approved the platform to transmit returns electronically. Authorization is specific to entity types: a platform approved for Individual returns may not be approved for Partnerships or Trusts. State authorization is separate from federal authorization and varies by jurisdiction. Ask for the vendor's current scope, entity types, and jurisdictions, and confirm whether a jurisdiction in marketing materials is live for transmission today. The program is public, so authorization is verifiable rather than a capability claim.
Q: Is AI tax preparation faster than traditional tax preparation?
A: The right question is not speed but where the time goes. AI preparation shifts the coordination and reconstruction work (document classification, workpaper population, open-item detection) from the reviewer's session to the preceding passes. Whether that produces faster delivery depends on return complexity, source document quality, and the firm's review model. The more useful question is: what does the reviewer find when they open the file, and how much of their session is reconstruction versus judgment?
Q: How do I evaluate an AI tax vendor's research capability?
A: Ask where the regulatory answers come from and whether those sources are current. A platform that grounds answers in the Internal Revenue Code, Treasury Regulations, and court cases, and surfaces the specific provision behind each answer, is a different product from one that generates responses from general training data. The citation is what distinguishes a researched position from a confident estimate. Ask the vendor to show a specific answer and trace it back to the provision it came from.

How AI Tax Return Review Works in Instead

How 1065 Filing Works With Instead
.png)
The 1040 Filing Workflow in Instead: Intake to E-File

Fixed Assets in Instead: How Depreciation Feeds the Return

.png)
