Reproducible where the evidence allows

How we test AI interview assistants

This methodology covers six desktop product reviews mapped into one July 2026 evidence matrix. It explains what was actually normalized, what we tested, what the outcome labels mean, what was not recorded consistently, and how to request a correction or retest.

Test families

What the review set examines

01

Interview and coding output

Behavioral, system-design, clean coding, debugging, and cluttered-screen scenarios check whether the first visible answer addresses the current task and preserves its constraints.

02

Context control

Area capture, selected text, full-screen capture, multi-window context, uploaded notes, and visible context buffers are checked where the tested product exposes them.

03

Desktop interaction

Focus or blur events, cursor changes, shortcut propagation, mouse interaction, screen-capture hiding, and receiver-side screen sharing are recorded separately.

04

System visibility

Normal user-visible surfaces such as Activity Monitor, Applications, and Windows Task Manager are checked when a product makes local invisibility or zero-trace claims.

05

Live usability

Audio capture, question detection, follow-ups, answer readability, controls, and second-screen workflows are evaluated only when the published test provides enough direct evidence.

06

Pricing and public claims

Public product pages, support documentation, and linked terms are recorded with the test date. Pricing can change and is not treated as a permanent product property.

Outcome definitions

The labels do not overreach the test

Passed

The expected behavior worked in the published test.

Mixed

The behavior worked only partly, varied across runs, or passed with an important limitation.

Failed in our test

The expected behavior did not work in the published test.

Not found

We did not find the feature in the tested plan or flow. This does not mean it is missing from other plans or versions.

Not tested

We did not test this behavior strongly enough to publish a result.

Normalization boundary

Shared labels, not an identical lab protocol

  • The criterion names, expected behaviors, and five outcome labels are normalized across the six reviews.
  • Test prompts, operating systems, builds, permissions, evidence layers, and repeat counts were not standardized retrospectively.
  • Passed, Mixed, and Failed cells link to direct published evidence. Not found cells may link to “Where we checked.” Not tested cells do not claim an evidence URL.
  • The matrix is a source-backed cross-review comparison. It is not a claim that every product completed the same controlled benchmark run.
Test inventory

Products, plans, versions, and scope

Product Platform Plan / flow Version Published limitation
Cluely macOS Free tier 2.1.19 Free-tier macOS flow on Cluely 2.1.19. Hardware model and OS build were not recorded in the published evidence.
Interview Coder macOS Free / limited desktop flow 3.0.1 Free / limited macOS desktop-flow test on Interview Coder 3.0.1. Paid tiers and later versions may behave differently.
LockedIn AI macOS Plan not recorded 1.7.7 LockedIn AI 1.7.7 was tested on macOS. Paid-plan state was not recorded; Duo was not evaluated.
ULTRACODE AI macOS + Windows Free demo / trial flow 8.10.0 ULTRACODE 8.10.0 was tested on the shown macOS and Windows setups. Future versions and other configurations may behave differently.
Parakeet AI macOS Free trial 3.9.3 Free-trial desktop test. Mobile operation was documented publicly but not tested; later versions and other audio routes may behave differently.
Final Round AI macOS + Windows 11 Free trial 2.5.0 Final Round AI 2.5.0 was tested on macOS and Windows 11. Scan Code and live-answer evidence came from macOS; focus and Task Manager checks came from Windows 11. Cursor behavior failed, while screen sharing was mixed because cursor changes and native tooltips could expose assistant interaction. Shortcut isolation, explicit follow-up controls, and phone operation were not tested.
Known limitations

What this benchmark cannot establish

  • It does not prove that a product always passes or fails across every version, plan, operating-system build, permission set, model, or customer configuration.
  • Hardware model, OS build, granted permissions, and repeat count were not recorded consistently across the first six reviews. Missing fields are disclosed instead of reconstructed.
  • “Not found” means the feature was not found in the tested plan or flow. It does not mean the feature cannot exist elsewhere.
  • Receiver-side screen-share evidence is stronger than an app's own capture preview. An own-preview result is not scored as a receiver-side result.
  • Public pricing and product claims are snapshots as of July 2026 and may change after publication.
Version history

Substantive benchmark changes

  1. Cluely macOS review published

    Free-tier coding, answer-shape, focus, cursor, shortcut, and local-visibility evidence.

  2. Interview Coder macOS review published

    Screen-share, focus, cursor, shortcut, audio, coding-context, and pricing evidence.

  3. LockedIn AI macOS review published

    Focus, shortcut, audio, process identity, auto-mode, answer-shape, and workflow evidence.

  4. ULTRACODE macOS and Windows review published

    Receiver-side screen-share, focus, cursor, shortcut, process, context, pricing, and macOS bundle evidence.

  5. Four-product evidence matrix and methodology published

    Added shared criteria and outcome definitions, limitations, methodology, organization attribution, and CSV/JSON downloads.

  6. Evidence audit and dataset license published

    Corrected audio, screen-share, shortcut, context, follow-up, and version claims; separated direct evidence from scope notes; and added a versioned data license.

  7. Parakeet AI macOS review added

    Added repeat Auto Answer context, screen-share, focus, cursor, shortcut, process-identity, workflow, pricing, and sanitized local-log security evidence to the five-product matrix.

  8. Final Round AI macOS and Windows review added

    Added Scan Code context, live-answer, focus, cursor, screen-share, Task Manager, pricing, and scoped Not tested outcomes to the six-product matrix.

CTRLpotato Benchmark Data License 1.0

Reuse the benchmark data with attribution

You may copy, analyze, quote, and redistribute the benchmark CSV and JSON, including commercially, as long as you attribute CTRLpotato and link back to the benchmark.

When citing the data, preserve the tested version, test date, and any relevant limitations so the findings stay in context.

This license covers only the cross-review evidence-matrix dataset (CSV and JSON). It does not apply to the website design, source code, logos, screenshots, videos, or third-party trademarks and content.

Corrections and retests

Send the version and reproducible setup

If a finding is outdated or you can reproduce a different result, send the product name, version, platform, plan, permissions, relevant settings, and the smallest sequence that reproduces it.

Request a correction or retest