You have a result you cannot trust enough to publish

That is the problem I solve. Outputs cannot be rebuilt, or reviewer concerns remain unexplained.

Publishing requires a record that holds under scrutiny. I document what matches, what does not, and what cannot be tested from the available materials. I work first with university PIs and their research teams.

I take a pipeline and produce something a reviewer, a collaborator, or a future reader can verify: a reproducible script, a transparent audit trail, and a written record of every decision.

PhD, Psychology (UNSW) · BA, Linguistics (CUHK) — I have first-hand experience across the research lifecycle.

What this covers

I cover quantitative research across EEG/ERP, eye-tracking, speech, corpus, psychometric, and systematic-review studies. Codebook construction and annotation. Pipeline documentation to the point where an independent researcher can rerun from the stated inputs. Discrepancy reporting with a full record of what matched, what did not, and what was not testable from the available data.

Methods and data that can be scoped

The appropriate deliverable depends on the data, design, validation standard, and what has already been demonstrated. The statuses below distinguish published-data reproduction from a tutorial or synthetic mechanism test.

If your research involves…A pipeline can cover…Current evidence
EEG / ERPFiltering, epochs, component measures, statistical checks, and documented handoverPublic EEG reproduction: five components cleared and two amber — details →
Eye-trackingFixation and reading-time measures, exclusions, visualisation, and statistical modellingTwo public-data examples: one complete-case corpus reproduction and one study with 20 of 21 reported results matched
Speech or phoneticsTranscription, alignment, acoustic measures, error review, and reportingCurrent end-to-end speech pipeline evidence is synthetic; natural-audio validation remains necessary
Child speechMeasurement and error-analysis workflows for a defined research taskSynthetic mechanism test only; no real children's voices and no screening or diagnostic validation
Corpus and text researchCorpus preparation, coding, counts, classifiers, and comparison with published analysesPublished corpus reproductions are available; the sentiment tutorial is narrowly calibrated to English data
Experiments, surveys, coding, reviews, and translationA study-specific workflow with explicit human checks and validation criteriaCurrent examples include a mixture of published-data reproductions, real-data tutorials, and small synthetic tests; status is stated per record

See the Evidence Library for the data-status badge, denominator, limitations, and comparison table for each example.

Where to start

Send a paper, a data link, or a bounded question. If the analysis is already documented but cannot be rerun, describe where it breaks. If there is no dataset yet, ask whether a planned pipeline is auditable: that is a valid starting point too. I take the submission and return a written statement of whether the work is in scope, what it would involve, and what it would cost. You do not need a call: everything happens in writing before any engagement begins.

Literature and synthesis

Turn a reading list into a structured matrix of designs, samples, measures, and reported findings.

Current example: information extracted from eleven real-paper abstracts. This is an extraction tutorial, not a completed systematic review. Demo →

Instruments and experiments

Check questionnaire wording, translation choices, materials, and preregistration logic against stated sources and decision rules.

Current example: GAD-7 and PHQ-9 items compared with published Chinese versions, including a deliberately inserted error. This demonstrates the checking workflow, not general translation accuracy. Demo →

Grant proposals

Map a draft against the published assessment criteria and identify where the rationale, design, feasibility, or evidence needs clarification.

Current example: a criterion-by-criterion alignment exercise using the 2026/27 GRF2 criteria. It is not a funder score or a prediction of funding. Demo →

From a pipeline that will not rerun

This path covers one of three named failure states: a script that errors on rerun, outputs that do not match the reported results, or a pipeline whose documentation is not enough for someone else to rerun it. The work is bounded to the stated pipeline, the stated dataset, and the stated outputs. It does not cover manuscript editing, statistical consulting on new study designs, or any analysis that requires data that cannot be shared.

Fit and delivery summary

I start with a written scope document and a fixed fee agreed before any work begins. Deliverables depend on scope, and they always include a record of what was run, what the outputs were, where they matched the reported results, and where they did not. Nothing is delivered verbally. The handover is a document someone who was not present for the work can read.

Engagement levels are compared on the Process & Pricing page. For administrative, teaching, or private-assistant work, see the FAQ. For an isolated AI research-team setup, see Own an AI Team.

Evidence from published-data audits

These records are published-data audits, not the full consulting offer. They show the checking method through stated comparisons and recorded outputs. Client engagements may instead build or repair a bounded research-analysis pipeline.

I re-run a published analysis from its available data and code, then provide a side-by-side comparison with the paper. The result may be an exact match, a close match, or a documented discrepancy; the comparison table shows which.

Published-data checkWhat was checkedResult
Public EEG reproductionSeven ERP components on public NEMAR dataFive cleared; two remain amber
L2 word-processing studyNineteen reported checksNineteen reproduced; four were close rather than exact
Metadiscourse corpus studyTwenty-four reported resultsTwenty exact; four discrepancies documented
Multimodal eye-tracking studyTwenty-one reported resultsTwenty matched; one mismatch documented
Public eye-tracking corpusFixation and reading-time measures on a stated complete-case subsetMeasures and effects reproduced for that subset

These are bounded audits, not claims that every analysis in each paper was rerun. Each Evidence Library record identifies the available data, the denominator used, the matched results, and the unresolved differences.

A trial analytics deliverable applies the same process to one bounded question from your paper: recover the relevant result where possible, trace any discrepancy, and hand over the documented workflow.

How the work is checked

I remain the accountable person for scoping, methodological decisions, review, and delivery. Specialist AI roles may support planning, analysis, drafting, and quality checks, but their outputs are reviewed against the source data, paper, code, or agreed criteria before handover.

The working sequence is: understand the data and research question → run basic integrity checks → build in an open, inspectable stack → compare the output with the paper or stated expectations → document the result, limitations, and handover.

If you prefer to operate the machinery yourself, an isolated AI research-team setup is available separately. Process & Pricing →

What is not in scope

Manuscript editing and writing. Statistical advice on new study designs. Production systems and live experiments. Urgent rescues with no named output to anchor the work. Any analysis whose data cannot be shared.

Contact

Email your paper, public data link, or bounded research question. No call is needed to start.

PhD, Psychology (UNSW) · BA, Linguistics (CUHK)

Pricing: One trial analytics deliverable on your topic starts from US$1,000 (approximately HK$8,000). Scope and acceptance criteria are agreed before work begins; if the completed deliverable does not meet them, the fee is waived. This assurance applies only to the trial analytics deliverable.

Larger engagements are quoted as fixed-fee milestones rather than hourly work.

Email

Replies within 24 hours on business days.

Working arrangement

Fully remote, with scheduling across time zones

Fixed-fee pricing — no hourly meter.