Spreadsheet assistants are not interchangeable. A grid-native assistant can act inside the workbook, while a file-analysis tool may run code or reason over an uploaded copy. That difference is documented workflow fit, not proof that one candidate writes more accurate formulas.
Candidates by documented workflow
| Candidate | Documented workflow reason to shortlist | Local test required |
|---|---|---|
| Microsoft 365 Copilot in Excel | Excel-native formula, table, chart, and workbook interactions described in Microsoft documentation | Exact desktop or web build, plan, formula correctness, reference integrity, and audit trail |
| Gemini in Google Sheets | Sheets-native assistance and in-grid functions described in Google Workspace documentation | Current edition, region, admin enablement, bulk operations, formulas, and source-row preservation |
| Claude for Excel or file analysis | Excel add-in and file workflows described in Anthropic documentation | Current plan and model, workbook coverage, dependency preservation, and error recovery |
| ChatGPT data analysis or Excel workflow | File-analysis and Excel workflows described in OpenAI documentation | Tools present in the account, calculation reproducibility, workbook round-trip, and data policy |
Check the provider's live plan, region, account, and admin settings. Do not assume a free or paid tier includes a particular model, file tool, add-in, code execution environment, or quota. Plan pages establish availability; they do not establish correctness on your workbook.
A blind local spreadsheet test
- Create representative workbooks. Include formulas, named ranges, multiple sheets, dates, blanks, errors, filtered rows, locale-specific separators, charts, pivots, and any macros or external links you use.
- Write expected answers first. Keep known totals, reference formulas, reconciliations, and deliberately planted errors. Use the same natural-language requests for every candidate.
- Blind output review where practical. Export formulas, analyses, or proposed changes without provider names and have a spreadsheet-literate reviewer score them.
- Score correctness and review cost. Check formula results, sheet and range references, row coverage, dependency preservation, repeatability, unsupported assumptions, time saved, and human correction time.
- Test the exact production surface. Record account plan, model, add-in version, platform, date, settings, and file size. Re-run after material product changes.
Failure risks to test, not assume
Test each generated formula on rows with known results and reconcile totals independently. Inspect every referenced sheet, range, named object, and assumption. For money, legal, regulatory, payroll, tax, or other consequential work, require a qualified human review and preserve the original workbook plus a change log.
Large row counts need their own protocol. Context-window size and file acceptance do not prove that every row was processed or that multi-sheet dependencies survived. Use row-count checks, hash or control totals, boundary cases, and known aggregates. When a file exceeds practical limits, move the computation to a controlled database or chunking workflow and verify the recombined result.
Calculate your cost →·Compare this model →·Find your model →