Reproducible where the evidence allows

How we test AI interview assistants

This methodology covers the four desktop products in the July 2026 benchmark. It explains what we tested, what the outcome labels mean, what wasn't recorded consistently, and how to request a correction or retest.

Test families

What the review set examines

01

Interview and coding output

Behavioral, system-design, clean coding, debugging, and cluttered-screen scenarios check whether the first visible answer addresses the current task and preserves its constraints.

02

Context control

Area capture, selected text, full-screen capture, multi-window context, uploaded notes, and visible context buffers are checked where the tested product exposes them.

03

Desktop interaction

Focus or blur events, cursor changes, shortcut propagation, mouse interaction, screen-capture hiding, and receiver-side screen sharing are recorded separately.

04

System visibility

Normal user-visible surfaces such as Activity Monitor, Applications, and Windows Task Manager are checked when a product makes local invisibility or zero-trace claims.

05

Live usability

Audio capture, question detection, follow-ups, answer readability, controls, and second-screen workflows are evaluated only when the published test provides enough direct evidence.

06

Pricing and public claims

Public product pages, support documentation, and linked terms are recorded with the test date. Pricing can change and is not treated as a permanent product property.

Outcome definitions

The labels do not overreach the test

Passed

The expected behavior worked in the published test.

Mixed

The behavior worked only partly, varied across runs, or passed with an important limitation.

Failed in our test

The expected behavior did not work in the published test.

Not found

We did not find the feature in the tested plan or flow. This does not mean it is missing from other plans or versions.

Not tested

We did not test this behavior strongly enough to publish a result.

Test inventory

Products, plans, versions, and scope

Product Platform Plan / flow Version Published limitation
Cluely macOS Free tier 2.1.19 Cluely 2.1.19 on the free-tier macOS flow. Hardware model and OS build were not consistently recorded.
Interview Coder macOS Free / limited desktop flow 3.0.1 Free / limited macOS desktop-flow test. Paid AI tiers and later versions may behave differently.
LockedIn AI macOS Desktop app with stealth settings 1.7.7 LockedIn AI 1.7.7 on macOS. Paid-plan state was not consistently recorded; Duo was not evaluated.
ULTRACODE AI macOS + Windows Free demo / trial flow 8.10.0 ULTRACODE AI 8.10.0 was tested on macOS and Windows; the published binary inspection covered the signed macOS v8.9.0 bundle.
Known limitations

What this benchmark cannot establish

  • It does not prove that a product always passes or fails across every version, plan, operating-system build, permission set, model, or customer configuration.
  • Hardware model, OS build, granted permissions, and repeat count were not recorded consistently across the first four reviews. Missing fields are disclosed instead of reconstructed.
  • “Not found” means the feature was not found in the tested plan or flow. It does not mean the feature cannot exist elsewhere.
  • Receiver-side screen-share evidence is stronger than an app's own capture preview. The matrix distinguishes those evidence levels.
  • Public pricing and product claims are snapshots as of July 2026 and may change after publication.
Version history

Substantive benchmark changes

  1. Cluely macOS review published

    Free-tier coding, answer-shape, focus, cursor, shortcut, and local-visibility evidence.

  2. Interview Coder macOS review published

    Screen-share, focus, cursor, shortcut, audio, coding-context, and pricing evidence.

  3. LockedIn AI macOS review published

    Focus, shortcut, audio, process identity, auto-mode, answer-shape, and workflow evidence.

  4. ULTRACODE macOS and Windows review published

    Receiver-side screen-share, focus, cursor, shortcut, process, context, pricing, and macOS bundle evidence.

  5. Four-product benchmark and methodology published

    Added normalized outcome definitions, limitations, methodology, organization attribution, and CSV/JSON downloads.

Corrections and retests

Send the version and reproducible setup

If a finding is outdated or you can reproduce a different result, send the product name, version, platform, plan, permissions, relevant settings, and the smallest sequence that reproduces it.

Request a correction or retest