How We Work — From Raw Data to Final Results

Nothing is hidden. Every engagement runs the same five steps, in one of two modes. Replication — a published study's numbers re-run from its public data (what the demos show, and what the trial does on your own published paper). Build — your own data, where there is no answer key, and verification has to be built in rather than checked against a paper. The steps below show both; the discipline is the same.

A PhD in psycholinguistics (UNSW), a BA in linguistics (CUHK), and years as a university researcher — these pipelines are built to work the way research actually happens.

  1. 1

    Know the data — public or yours

    What happens: Replication — the public deposit behind the paper (OSF, PLOS, journal supplementary files, or GitHub) is located and verified. Build — your dataset is received and inventoried: structure, codebook, provenance, and where every variable came from.

    Why it matters: A fixed, documented starting point means stable results and nothing hand-picked. On new data, this inventory becomes the project's codebook — the record that keeps the analysis defensible.

  2. 2

    Load and sanity-check

    What happens: Row and column counts, value ranges, and missingness appear on screen before any analysis begins.

    Why it matters: The first question any reviewer asks is whether the data really is what it claims to be. We answer it visibly, up front.

  3. 3

    Build or rebuild the analysis

    What happens: Replication — the paper's core models are rebuilt from scratch in an open statistical stack, not replayed from the original software. Build — the analysis is constructed from your design and codebook in the same open stack.

    Why it matters: Independence is the point: we reconstruct the logic rather than echo an output — and the stack is open enough for any reviewer to follow.

  4. 4

    Verify — against the paper, or against the data itself

    What happens: Replication — published statistics versus the re-run, item by item, in a single comparison table. Build — a verification battery: diagnostics, sensitivity and robustness checks, recovery of known effects — because with new data there is no answer key.

    Why it matters: This is the proof. Where numbers match, you see it exactly; where they differ, the difference is flagged in the report — never explained away. New data gets the same discipline, with the verification battery standing in for the comparison table.

  5. 5

    Report and hand over

    What happens: A comparison table (or verification report), a plain-language summary of what reproduced and what didn't — plus the reproducible script and data.

    Why it matters: The deliverable is yours to keep — your team can re-run everything in your own lab, after the engagement ends.

Every pipeline follows the same discipline — load → check → analyse → validate → report. What you see in the demo is exactly what you get on your own data; your data just gets its own verification battery instead of a published answer key.

Want to see the whole machine, including the code? The fastest route is the trial — every script, documented, run on your own published data, in a matter of days. See the FAQ →

This page covers the research method — the admin agent and teaching modernization each have their own process, also in the FAQ.