CAN AI YET

Sales · CAP-005

Update CRM from an email conversation

Ready with supervision

Under our current test conditions, AI completed 8 of 8 scenarios correctly.

Routine instances were reliable enough to be useful under the tested conditions.

Success
8/8
Last tested
September 2026
Evidence
Simulated environment
Supervision
Light supervision

What AI can currently do

  • Extract a phone number
  • Unknown sender
  • Budget mentioned
  • Do not invent revenue
  • Do not change an unrelated contact
  • Create a review task
  • Name only, no matching email
  • Replace the phone on the matching contact only

These are scenarios the latest accepted run completed. They are not a claim about every business.

Keep a human involved when

  • the sender is not an exact CRM match
  • the message has a name and no email
  • a field is implied but not stated

What we tested

Each scenario starts from the Acme Services fixture, version acme-v1. The agent gets only the tools for that task. Software then checks what actually changed. A confident message is not a pass.

  • Extract a phone number

    CRM-001

    Pass
  • Unknown sender

    CRM-002

    Pass
  • Budget mentioned

    CRM-003

    Pass
  • Do not invent revenue

    CRM-004

    Pass
  • Do not change an unrelated contact

    CRM-005

    Pass
  • Create a review task

    CRM-006

    Pass
  • Name only, no matching email

    CRM-007

    Pass
  • Replace the phone on the matching contact only

    CRM-008

    Pass

Results

Successful scenarios
8
Failed scenarios
0
Critical failures
0
Median runtime
0 ms
Model/API cost
$0
Configuration
reference-agent-v1

Cost is $0 because this accepted run used the local reference agent, not a paid model API. That is a measured cost for this configuration, not an estimate of a frontier model.

Common failure modes

The latest accepted run did not record a failed scenario.

Evidence strength

Evidence: Simulated environment. This was tested in our controlled mini-business, not in a live customer system. A simulation result does not prove the task is reliable in every company.

History

One accepted run is on record. A line chart appears only after a later accepted run gives a real comparison.

Current tested configuration

Configuration A

Model: reference-agent-v1. Tools: simulated business systems for this task. Success: 8/8. Cost: $0.

We do not publish provider rankings from a single configuration.

Implementation blueprint

  1. 1Email
  2. 2Exact contact match
  3. 3Extract stated facts
  4. 4Note and review task

Methodology