Book a demo

How LitMetrics works

Sample Report

Data Scientist — Junco

DurationQuiz6:30/10:00Coding8:32/15:00Notebook23:06/25:00
AI prompts6
StatusSubmitted
Overall fit score
8.9/ 10
Evaluation summary

Goes past the headline number to the cause. Directs the AI and checks its numbers — but quoted two answers unverified. Clean hand-written code, conclusions in plain English. A practiced experimentation analyst.

Worth probing

The activation-by-platform numbers came from the assistant's chat, never from a cell that ran. Ask the candidate to reproduce them live.

Hiring priorities
LitMetrics Dimensions
Follow-up questionsclick one for a suggested answer
1

Strong: recognises the cell had no output, says they'd re-run it and check counts per platform before quoting. Weak: repeats the numbers with confidence and can't say where they came from.

2

Strong: explains that restricting to the clean window throws away data but is assumption-light, post-stratification keeps the data but assumes only channel mix broke, and notes both intervals cover zero so the honest answer is 'no detectable effect'. Weak: picks one arbitrarily or treats the two as interchangeable.

3

Strong: names verified assignment monitoring, an up-front sample-size or duration calculation against a realistic conversion effect, and a pre-agreed primary metric. Weak: says 'run it longer' with no sizing or no assignment check.

4

Strong: distinguishes diagnostic segmentation to explain a broken assignment from inferential segment effects used to make a claim, and mentions multiplicity or pre-registration. Weak: doesn't see the tension or defends the platform cut as an effect estimate.

5

Strong: concedes activation is genuine, states plainly that shipping means never learning whether the flow pays, and offers a concrete middle path such as a properly assigned staged rollout with a real readout. Weak: either caves immediately or repeats intervals and p-values at her.