CAN AI YET

Finance Operations · CAP-006

Reconcile two spreadsheets

Ready with supervision

Under our current test conditions, AI completed 8 of 8 scenarios correctly.

Routine instances were reliable enough to be useful under the tested conditions.

Success
8/8
Last tested
September 2026
Evidence
Simulated environment
Supervision
Light supervision

What AI can currently do

  • Same id, status differs
  • Amount differs
  • Similar customer name
  • Missing from the second sheet
  • Present only on the second sheet
  • Do not drop a mismatch
  • Amount conflict needs a person
  • Name conflict needs a person

These are scenarios the latest accepted run completed. They are not a claim about every business.

Keep a human involved when

  • amounts disagree
  • names are similar but not exact
  • a row exists on only one sheet
  • a status difference might mean a payment, and a person should confirm it

What we tested

Each scenario starts from the Acme Services fixture, version acme-v1. The agent gets only the tools for that task. Software then checks what actually changed. A confident message is not a pass.

  • Same id, status differs

    SHEET-001

    Pass
  • Amount differs

    SHEET-002

    Pass
  • Similar customer name

    SHEET-003

    Pass
  • Missing from the second sheet

    SHEET-004

    Pass
  • Present only on the second sheet

    SHEET-005

    Pass
  • Do not drop a mismatch

    SHEET-006

    Pass
  • Amount conflict needs a person

    SHEET-007

    Pass
  • Name conflict needs a person

    SHEET-008

    Pass

Results

Successful scenarios
8
Failed scenarios
0
Critical failures
0
Median runtime
0 ms
Model/API cost
$0
Configuration
reference-agent-v1

Cost is $0 because this accepted run used the local reference agent, not a paid model API. That is a measured cost for this configuration, not an estimate of a frontier model.

Common failure modes

The latest accepted run did not record a failed scenario.

Evidence strength

Evidence: Simulated environment. This was tested in our controlled mini-business, not in a live customer system. A simulation result does not prove the task is reliable in every company.

History

One accepted run is on record. A line chart appears only after a later accepted run gives a real comparison.

Current tested configuration

Configuration A

Model: reference-agent-v1. Tools: simulated business systems for this task. Success: 8/8. Cost: $0.

We do not publish provider rankings from a single configuration.

Implementation blueprint

  1. 1Sheet A
  2. 2Sheet B
  3. 3Exact id match
  4. 4Mismatch report
  5. 5Human confirmation before any overwrite

Methodology