PhD, Psychology (UNSW) · BA, Linguistics (CUHK) — I have first-hand experience across the research lifecycle.
What this covers
I cover quantitative research across EEG/ERP, eye-tracking, speech, corpus, psychometric, and systematic-review studies. Codebook construction and annotation. Pipeline documentation to the point where an independent researcher can rerun from the stated inputs. Discrepancy reporting with a full record of what matched, what did not, and what was not testable from the available data.
Methods and data that can be scoped
The appropriate deliverable depends on the data, design, validation standard, and what has already been demonstrated. The statuses below distinguish published-data reproduction from a tutorial or synthetic mechanism test.
| If your research involves… | A pipeline can cover… | Current evidence |
|---|---|---|
| EEG / ERP | Filtering, epochs, component measures, statistical checks, and documented handover | Public EEG reproduction: five components cleared and two amber — details → |
| Eye-tracking | Fixation and reading-time measures, exclusions, visualisation, and statistical modelling | Two public-data examples: one complete-case corpus reproduction and one study with 20 of 21 reported results matched |
| Speech or phonetics | Transcription, alignment, acoustic measures, error review, and reporting | Current end-to-end speech pipeline evidence is synthetic; natural-audio validation remains necessary |
| Child speech | Measurement and error-analysis workflows for a defined research task | Synthetic mechanism test only; no real children's voices and no screening or diagnostic validation |
| Corpus and text research | Corpus preparation, coding, counts, classifiers, and comparison with published analyses | Published corpus reproductions are available; the sentiment tutorial is narrowly calibrated to English data |
| Experiments, surveys, coding, reviews, and translation | A study-specific workflow with explicit human checks and validation criteria | Current examples include a mixture of published-data reproductions, real-data tutorials, and small synthetic tests; status is stated per record |
See the Evidence Library for the data-status badge, denominator, limitations, and comparison table for each example.
Where to start
Send a paper, a data link, or a bounded question. If the analysis is already documented but cannot be rerun, describe where it breaks. If there is no dataset yet, ask whether a planned pipeline is auditable: that is a valid starting point too. I take the submission and return a written statement of whether the work is in scope, what it would involve, and what it would cost. You do not need a call: everything happens in writing before any engagement begins.
Literature and synthesis
Turn a reading list into a structured matrix of designs, samples, measures, and reported findings.
Current example: information extracted from eleven real-paper abstracts. This is an extraction tutorial, not a completed systematic review. Demo →
Instruments and experiments
Check questionnaire wording, translation choices, materials, and preregistration logic against stated sources and decision rules.
Current example: GAD-7 and PHQ-9 items compared with published Chinese versions, including a deliberately inserted error. This demonstrates the checking workflow, not general translation accuracy. Demo →
Grant proposals
Map a draft against the published assessment criteria and identify where the rationale, design, feasibility, or evidence needs clarification.
Current example: a criterion-by-criterion alignment exercise using the 2026/27 GRF2 criteria. It is not a funder score or a prediction of funding. Demo →
From a pipeline that will not rerun
This path covers one of three named failure states: a script that errors on rerun, outputs that do not match the reported results, or a pipeline whose documentation is not enough for someone else to rerun it. The work is bounded to the stated pipeline, the stated dataset, and the stated outputs. It does not cover manuscript editing, statistical consulting on new study designs, or any analysis that requires data that cannot be shared.
Fit and delivery summary
I start with a written scope document and a fixed fee agreed before any work begins. Deliverables depend on scope, and they always include a record of what was run, what the outputs were, where they matched the reported results, and where they did not. Nothing is delivered verbally. The handover is a document someone who was not present for the work can read.
Engagement levels are compared on the Process & Pricing page. For administrative, teaching, or private-assistant work, see the FAQ. For an isolated AI research-team setup, see Own an AI Team.
Evidence from published-data audits
These records are published-data audits, not the full consulting offer. They show the checking method through stated comparisons and recorded outputs. Client engagements may instead build or repair a bounded research-analysis pipeline.
I re-run a published analysis from its available data and code, then provide a side-by-side comparison with the paper. The result may be an exact match, a close match, or a documented discrepancy; the comparison table shows which.
| Published-data check | What was checked | Result |
|---|---|---|
| Public EEG reproduction | Seven ERP components on public NEMAR data | Five cleared; two remain amber |
| L2 word-processing study | Nineteen reported checks | Nineteen reproduced; four were close rather than exact |
| Metadiscourse corpus study | Twenty-four reported results | Twenty exact; four discrepancies documented |
| Multimodal eye-tracking study | Twenty-one reported results | Twenty matched; one mismatch documented |
| Public eye-tracking corpus | Fixation and reading-time measures on a stated complete-case subset | Measures and effects reproduced for that subset |
These are bounded audits, not claims that every analysis in each paper was rerun. Each Evidence Library record identifies the available data, the denominator used, the matched results, and the unresolved differences.
A trial analytics deliverable applies the same process to one bounded question from your paper: recover the relevant result where possible, trace any discrepancy, and hand over the documented workflow.
How the work is checked
I remain the accountable person for scoping, methodological decisions, review, and delivery. Specialist AI roles may support planning, analysis, drafting, and quality checks, but their outputs are reviewed against the source data, paper, code, or agreed criteria before handover.
The working sequence is: understand the data and research question → run basic integrity checks → build in an open, inspectable stack → compare the output with the paper or stated expectations → document the result, limitations, and handover.
If you prefer to operate the machinery yourself, an isolated AI research-team setup is available separately. Process & Pricing →
What is not in scope
Manuscript editing and writing. Statistical advice on new study designs. Production systems and live experiments. Urgent rescues with no named output to anchor the work. Any analysis whose data cannot be shared.
Contact
Email your paper, public data link, or bounded research question. No call is needed to start.
PhD, Psychology (UNSW) · BA, Linguistics (CUHK)
Pricing: One trial analytics deliverable on your topic starts from US$1,000 (approximately HK$8,000). Scope and acceptance criteria are agreed before work begins; if the completed deliverable does not meet them, the fee is waived. This assurance applies only to the trial analytics deliverable.
Larger engagements are quoted as fixed-fee milestones rather than hourly work.
Replies within 24 hours on business days.
Working arrangement
Fixed-fee pricing — no hourly meter.