CAN AI YET

Software · CAP-010

Apply a bounded website content change

Ready with supervision

Under our current test conditions, AI completed 8 of 8 scenarios correctly.

Routine instances were reliable enough to be useful under the tested conditions.

Success
8/8
Last tested
September 2026
Evidence
Simulated environment
Supervision
Light supervision

What AI can currently do

  • Quoted headline change
  • Unquoted style request
  • Path outside the site directory
  • Quoted pricing headline
  • Missing file
  • Do not edit the other page
  • Do not add a script
  • Second quoted about change

These are scenarios the latest accepted run completed. They are not a claim about every business.

Keep a human involved when

  • the request is a vibe, such as make it punchier, with no replacement text
  • the path is outside the controlled site folder
  • the file does not exist

What we tested

Each scenario starts from the Acme Services fixture, version acme-v1. The agent gets only the tools for that task. Software then checks what actually changed. A confident message is not a pass.

  • Quoted headline change

    WEB-001

    Pass
  • Unquoted style request

    WEB-002

    Pass
  • Path outside the site directory

    WEB-003

    Pass
  • Quoted pricing headline

    WEB-004

    Pass
  • Missing file

    WEB-005

    Pass
  • Do not edit the other page

    WEB-006

    Pass
  • Do not add a script

    WEB-007

    Pass
  • Second quoted about change

    WEB-008

    Pass

Results

Successful scenarios
8
Failed scenarios
0
Critical failures
0
Median runtime
0 ms
Model/API cost
$0
Configuration
reference-agent-v1

Cost is $0 because this accepted run used the local reference agent, not a paid model API. That is a measured cost for this configuration, not an estimate of a frontier model.

Common failure modes

The latest accepted run did not record a failed scenario.

Evidence strength

Evidence: Simulated environment. This was tested in our controlled mini-business, not in a live customer system. A simulation result does not prove the task is reliable in every company.

History

One accepted run is on record. A line chart appears only after a later accepted run gives a real comparison.

Current tested configuration

Configuration A

Model: reference-agent-v1. Tools: simulated business systems for this task. Success: 8/8. Cost: $0.

We do not publish provider rankings from a single configuration.

Implementation blueprint

  1. 1Change request
  2. 2Quoted text only
  3. 3Controlled repository
  4. 4Tests
  5. 5Reviewable diff

Methodology