Browse 45 working research pipelines: real-data reproductions, tutorials, simulated or synthetic mechanism tests. Each badge tells you what kind of evidence you are seeing; each walkthrough shows the method and its limitations.
This library is designed for audit before commitment: inspect what runs, what data support it, and where its limits are before deciding whether an approach fits your research.
Recommended starting points are ordered for visitor utility across university and research needs.
The order considers evidence readiness, first-screen clarity, consulting breadth, distinctiveness, risk and journey fit.
It is not a ranking of scientific importance, research quality or academic strength.
Choose a field, a library classification or a search term. Categories indicate each demo's primary route; some methods may be relevant across fields.
Library classification controls this page's filter. Demo badge reproduces the designation shown in the demo; the two fields are not a one-to-one scale.
Find your field, check the evidence badge, read the walkthrough and limitations, then run the demo. If the approach looks relevant, bring the paper, dataset or research question you want to test.
Badges distinguish real-data reproductions from tutorial, simulated, and synthetic work. They are evidence labels, not quality scores.
Field
Keyword filters
Library classification
Showing 45 of 45 demos
5 effects cleared; 2 remain amber.
Follow a transparent public-data EEG pipeline from raw signals to group statistics, with five effects clearing the pre-specified gate and two retained as amber findings.
What it does: Reproduces classic event-related potential effects from a public multi-paradigm EEG resource, with pre-specified gates and honest amber labels.
Demo badge: published
Headline result: Five of seven pre-specified component effects cleared the reproduction gate; two remained sub-threshold but directionally consistent.
Start here: Follow the walkthrough from raw download through cleaning, then compare each component result against the published benchmark.
The data:
ERP CORE (Kappenman et al., 2021, Scientific Data), the canonical open EEG resource: 40 participants across six paradigms (~133,000 single trials) covering seven classic components — N170, MMN, N2pc, N400, P3, LRP, ERN. Raw BIDS files were downloaded from the NEMAR mirror (nm000132 v1.1.1), SHA-256 verified, and processed with MNE-Python (automatic ICA, global 200 µV artifact ceiling, pre-specified measurement windows per component). All numbers on the dashboard are the locked outputs of that pipeline.
What you'd see, step by step:
Try or inspect it:
Follow the seven-step walkthrough, compare each step with the previous one, and inspect or download the supporting tables, trial-level CSVs, and scripts.
Limitations:
Artifact rejection differs from the original paper (automatic ICA + global 200 µV ceiling vs per-subject tailored thresholds), so participant counts differ slightly from Kappenman et al. Table 1 (N=34–39). LRP and ERN did not clear the pre-registered effect-size gate (amber, reported honestly — not a pipeline failure). The N400 cross-check matches the published anchor in sign and magnitude (Δdz=0.022) but falls outside a very tight 0.01 tolerance, which is documented rather than suppressed. No item random effects.
20 of 20 key tests matched; the reproduction is non-diagnostic.
Re-run the deposited acoustic data and compare the reported group-level patterns, while keeping the small sample and omitted post-hoc tests visible.
What it does: Re-runs acoustic analyses from a Cantonese child speech study and compares group prosody patterns, for readers checking published reproduction claims.
Demo badge: published
Headline result: Children with autism showed smaller pitch range on focused syllables than typical peers; twenty of twenty key tests matched the paper.
Start here: Inspect the deposited measurements, then follow cleaning and modelling through to the group comparison summary.
The data:
The dataset comes from a 2024 PLOS ONE study by Chen, Zhang, Zhou, Chan and colleagues, and contains acoustic speech measurements from 46 Cantonese-speaking children (23 with autism, 23 typically developing) across nearly 10,000 recorded sentences. The data and analysis code were deposited publicly by the original authors, so everything shown is real, peer-reviewed research material. Because the pipeline re-runs on the deposited snapshot, results are stable — but small differences from the published numbers are expected and are explained honestly during the demo.
What you'd see, step by step:
Try or inspect it:
Inspect the group-comparison table and prosody plots for the bundled study data. Contact the team for another run under an agreed data-handling arrangement.
Limitations:
The study is small by clinical standards (23 children per group), so effect sizes should be treated as indicative rather than definitive. The pipeline reproduces all directions of difference and statistical significance, but the raw test-statistic numbers differ from the paper's tables due to a difference in how model comparisons were sequenced — the conclusions are identical either way. Post-hoc pairwise comparisons are not re-run here, so fine-grained contrast tables from the original paper are taken on trust.
5 measures matched within 1 ms on a complete-case subset.
Check five published eye-movement measures and two reading effects in the Provo Corpus complete-case subset.
What it does: Recomputes five Provo Corpus eye-movement measures and two reading effects, for readers verifying published word-reading statistics.
Demo badge: published
Headline result: Five eye-movement measures matched published values at Pearson r equals 1.0000 within one millisecond.
Start here: Review the complete-case subset size, then follow measure computation through predictability and word-length contrasts.
The data:
This demo analyzes 153,566 of 230,412 word observations from 84 native English readers in the public Provo Corpus. Five eye-movement measures match the authors' published values at Pearson r = 1.0000 (within 1 ms); predictability effects are reproduced for first-fixation duration (15.4 ms, t(83) = 11.37, p = 1.3e-18, dz = 1.24) and gaze duration (43.1 ms, t = 15.24, p = 8.9e-26, dz = 1.66), alongside word-length effects. Limitation: only 153,566 of 230,412 word observations are analyzed, so this is not a full-observation reanalysis.
What you'd see, step by step:
Try or inspect it:
Open https://demo.patrickchu.net/eetracking-real/ and run the pipeline live — every figure, table and statistic can also be exported.
Limitations:
Only 153,566 of 230,412 word observations carry complete data across all five measures and are analyzed (the rest are excluded per the published definitions, e.g. first-pass skips); this is not a full-observation reanalysis. Word-frequency effects are not included because the official Provo files publish no word-frequency column — only predictability and word-length effects are claimed.
20 of 24 results matched exactly; 4 were close.
Re-run the archived code and corpus to inspect 20 exact matches and four documented discrepancies.
What it does: Re-runs corpus comparisons between translated and native English across four genres, for readers checking twenty-four published statistical results.
Demo badge: published
Headline result: Twenty of twenty-four published results matched exactly; four were close with documented reasons.
Start here: Review the thousand-text design across provenance groups and genres, then open the comparison table beside published values.
The data:
The dataset is 1,000 academic and general-interest texts — 500 pieces of translated English and 500 pieces of native English — drawn from two established public corpora and spanning four genres (academic, news, prose, fiction). The data and the authors' original analysis code are openly archived by Chou, Li, and Liu (PLOS ONE, 2023) and can be downloaded by anyone. Because the dataset is fixed and archived, results are stable; the numbers you see in the demo will be the same every time.
What you'd see, step by step:
Try or inspect it:
Inspect the displayed results table and compare translated with native English by genre. For a run on other material, contact the team to agree data handling first.
Limitations:
The corpus is real but modest in size (1,000 texts, four genres), so findings describe these specific collections, not all English writing. Four of the 24 replicated results are close but not exact — two because the original paper used different statistical settings than its own archived code, and two because of what appears to be a misprint in the published table. Effect sizes throughout are small, meaning the translated/native differences are real but not large in practical terms. Any live re-run will produce the same numbers as the demo, since the data is static.
Design-field agreement ranged from about 0.76 to 0.88.
Extract structured fields from 11 pinned abstracts and inspect agreement, omissions and the limits of a small non-exhaustive corpus.
What it does: Extracts structured fields from eleven pinned generative-AI-in-education abstracts, for researchers learning literature-matrix workflows on a small corpus.
Demo badge: published
Headline result: Design-field agreement with a keyword baseline ranged from about 0.76 to 0.88 across the pinned abstract set.
Start here: Assess the eleven-abstract scope, then watch each abstract populate a matrix row with cross-checked design labels.
The data:
The demo uses 11 real, open-access, top-cited papers on ChatGPT in higher education (2023–2026), fetched from OpenAlex and pinned for reproducibility. The sample is small and selective (a top-cited sample, not an exhaustive search), and because live data can drift, results are reported as ranges rather than single numbers.
What you'd see, step by step:
Try or inspect it:
Inspect the abstract-extraction matrix and compare two bundled papers. For a run on another literature set, contact the team first under an agreed data-handling arrangement.
Limitations:
Extraction is from abstracts only, not full texts, so some details are thinner. The LLM is stochastic but pinned to a cache for reproducibility; observed agreement with the baseline ranges from 0.76 to 0.88. The corpus is a small, top-cited sample, not an exhaustive review.
Re-run the deposited analyses and inspect the 20 matched results alongside the one borderline mismatch.
What it does: Re-runs deposited eye-tracking analyses comparing translation and paraphrase reading, for readers checking twenty-one published statistical verdicts.
Demo badge: published
Headline result: Twenty of twenty-one statistical verdicts matched the paper; eye behaviour differed between translation and paraphrase conditions.
Start here: Inspect which participant-level measures are available, then follow cleaning through models to the side-by-side results table.
The data:
The data come from a real published experiment (Ma, Han & Li, 2022, PLOS ONE) in which participants read translated and paraphrased texts while their eye movements were recorded; the original spreadsheets and analysis scripts are freely available online (OSF). Because the dataset is fixed and not live, results are stable and directly comparable to what appears in the published paper. There is no synthetic or simulated data involved.
What you'd see, step by step:
Try or inspect it:
Inspect the displayed results table, choose one of the 21 eye-tracking results, and compare the published and checked values. For another dataset, contact the team to agree data handling first.
Limitations:
The dataset is small (one published experiment, a limited number of participants and texts), so findings should not be generalised beyond the original study's scope. One result out of 21 does not match the published value — it sits right on the borderline of significance and is sensitive to minor differences between the original software and the replication software; this is flagged honestly rather than glossed over. An ordinal model used in the original paper for translation-difficulty ratings could not be re-run and remains a known gap.
Compare the archived experiments with the paper's core age-perception results, while showing the omitted sub-models and effect-size convention.
What it does: Compares two archived face-perception experiments with the paper's core age estimates, for readers assessing how far the numbers hold up.
Demo badge: published
Headline result: Twenty published statistics reproduced with matching direction and significance across both experiments.
Start here: Compare the two experimental designs, then step through estimates, tests, and the side-by-side number table.
The data:
The dataset comes from two experiments by Ji, Liao & Hayward (BMC Psychology, 2025) and is publicly archived on OSF alongside the original analysis scripts. Experiment 1 tested 88 young Hong Kong participants estimating the ages of White faces showing neutral, happy, or angry expressions; Experiment 2 added older participants and Asian face stimuli. The data are real and stable (a fixed, cleaned deposit), so results are consistent across runs rather than shifting with live feeds.
What you'd see, step by step:
Try or inspect it:
Ask the demo to show only Experiment 2 results — type 'exp2' when prompted — to see whether the older-participant effect still appears in isolation.
Limitations:
Both experiments used White or Asian face stimuli rated by Hong Kong participants, so findings may not generalise broadly. Sample sizes are modest (88 young adults in Exp 1; roughly 72 across age groups in Exp 2). One effect-size number differs from the paper because the paper and this pipeline follow different but equally valid conventions for calculating it — the underlying finding is the same. The pipeline reproduces the core paired comparisons and the participant-age modulation, but not every sub-model from the original R script.
Compare AI scores and feedback with two human ratings on 25 public US middle-school essays, while examining the validation and oversight required before local use.
What it does: Compares AI essay scores and feedback with two human raters on twenty-five public middle-school essays, for educators exploring rubric alignment limits.
Demo badge: real
Headline result: AI scores matched human ratings exactly sixty-eight percent of the time, similar to the human–human agreement rate.
Start here: Select an included anonymised essay, then compare both human scores and AI feedback against the displayed rubric.
The data:
The demo uses 25 real student essays drawn from a large public dataset of grade 7–8 persuasive writing, each independently scored by two trained human raters on a 1–6 scale. The data is real, not synthetic, but it comes from a US classroom context — results on Hong Kong student writing or different rubrics would need separate validation before any conclusions are drawn.
What you'd see, step by step:
Try or inspect it:
Select an included essay, compare both human ratings with the AI output, and critique the feedback against the displayed rubric. Do not upload student work.
Limitations:
The 25-essay sample is small, so treat the agreement figures as indicative ranges rather than fixed numbers — a larger local sample would tighten them considerably. All essays here are from US middle-school students writing in English; the AI's scoring reliability on Hong Kong student writing, or against a different institutional rubric, is unknown and would need its own validation study (typically 50–100 essays with two local raters). The AI also produces only a single overall score in this demo; separate scores for content, organisation, and language conventions are possible but not shown here.
Compare glossary-assisted and unassisted translations on 14 synthetic government-style sentences, with terminology hits separated from AI-judged quality.
What it does: Compares glossary-assisted and unassisted translation on fourteen synthetic government-style sentences, for teams testing terminology-memory gains.
Demo badge: synthetic
Headline result: Terminology accuracy reached roughly eighty to ninety percent without a glossary and one hundred percent with the twenty-two-term glossary.
Start here: Pick a bundled sentence, then compare side-by-side translations with and without the injected glossary.
The data:
14 Chinese-language government press-release sentences, hand-authored to include a mix of everyday official terms (e.g. 教育局, 立法會) and rarer policy-specific terms (e.g. 簡約公屋 → Light Public Housing, 明日大嶼 → Lantau Tomorrow Vision). The data is small and synthetic — it is a mechanism demo, not a full study. Because the underlying AI model can be updated at any time, exact numbers may shift slightly between runs; the demo quotes ranges from repeated runs, not single best scores.
What you'd see, step by step:
Try or inspect it:
Select another bundled government-style sentence and compare the glossary-assisted and unassisted outputs.
Limitations:
Only 14 sentences — enough to demonstrate the mechanism, not to publish. A real study would need 100+ sentences with professional translators setting the gold standard. The quality scores (fluency, adequacy) are generated by an AI judge, not a human, and both translation versions are already fluent, so those scores are close and should not be over-interpreted. Results are ranges across runs, not guaranteed single figures.
Inspect how an AI applies stated inclusion criteria to 16 hand-written abstracts, with every provisional decision kept open to human review.
What it does: Screens sixteen hand-written abstracts against stated inclusion criteria, for review teams testing AI-assisted abstract triage on a pilot set.
Demo badge: synthetic
Headline result: The screener agreed with gold labels on fifteen of sixteen bundled abstracts in the fixed pilot set.
Start here: Read the review question and inclusion criteria, then follow each abstract through verdict and reason to the summary panel.
The data:
Sixteen short research abstracts were written by hand for this demo: eight that genuinely fit the review question (technology-enhanced feedback on school students' writing) and eight that do not, covering common exclusion reasons such as wrong age group, wrong subject, or no real study data. The dataset is synthetic and intentionally clean; real-world records are messier, and a proper validation would need 100+ abstracts checked by two independent human raters. Because this is a fixed pilot set, the headline numbers are illustrative ranges rather than a certified benchmark.
What you'd see, step by step:
Try or inspect it:
Choose one of the supplied synthetic abstracts, inspect the provisional include/exclude decision, and test which PICOS criterion drove it.
Limitations:
This is a 16-abstract pilot built on hand-written, unusually tidy records — not a real-world benchmark. Gold labels come from a single rater, not the two independent reviewers a real review requires. The one missed study (a meta-analysis mis-labelled by the demo) shows the AI is not error-free; recall figures in production will likely be lower and should be reported as a range. A live deployment would also need deduplication and full-text screening stages not shown here.
Compare three AI-generated drafts with published Chinese items to inspect the workflow—not to validate a finished bilingual instrument.
What it does: Walks through back-translation quality checks on eleven bilingual questionnaire items, for researchers testing translation workflow before field work.
Demo badge: illustrative
What this demo produces: The workflow ranks which of three Chinese drafts best preserves meaning against the official published version.
Start here: Open the English items and three Chinese drafts, then follow forward translation and back-translation scoring.
The data:
The demo uses 11 public-domain items from the GAD-7 and PHQ-9 scales. It compares three translation drafts (plain, glossary-guided, and colloquial Cantonese) against the official Traditional Chinese versions. The data is synthetic in the sense that the drafts are generated by an AI, and results may vary slightly with each run.
What you'd see, step by step:
Try or inspect it:
Compare the plain, glossary-guided, and colloquial drafts for any of the 11 bundled items. Contact the team for another run under an agreed data-handling arrangement.
Limitations:
The AI acts as both translator and judge, so this validates the workflow, not a finished instrument. Results are ranges, not exact numbers, because the AI is stochastic. The comparison with published Chinese is indicative, as regional variants exist.
A live pipeline checks the archived survey data against the paper—reproducing the available results while showing why the headline before/after analysis cannot be verified.
What it does: Checks a deposited COVID-19 survey against its published analyses, showing which results reproduce and which cannot be verified from the archive.
Demo badge: published
Headline result: Twenty of twenty-one checkable items reproduced; worry correlations and action rates matched published figures exactly.
Start here: Inspect which variables are in the public deposit, then compare summary statistics and correlations with the paper.
The data:
A publicly archived survey of 3,032 adults in Hong Kong, Singapore, and the United States, collected around the WHO's March 2020 pandemic declaration (published data deposit by Prof Catherine Wing-Man Yeung, CUHK Business School, and co-authors). The dataset is real, publicly available, and fixed — it will not change between demos, so results are stable. Because the deposit does not include the date each person responded, the before/after comparison from the paper cannot be fully re-run; results for that part are therefore reported as ranges, not single figures.
What you'd see, step by step:
Try or inspect it:
Inspect the deposit-coverage table to see which analyses can and cannot be reproduced. Contact the team for another run under an agreed data-handling arrangement.
Limitations:
The dataset is real but the deposit is incomplete: response dates were not archived, so the paper's headline before/after ANOVAs cannot be independently verified. Worry scores and action rates are therefore shown as overall figures that fall between the published before and after values, not as a true replication of the time-split analysis. All other results — correlations, sample counts, action rates — reproduce cleanly. The fix is straightforward: if the authors share even a simple before/after flag, the remaining checks run in minutes.
Re-run the deposited data and scripts to inspect 15 reproduced comparisons and four transparently labelled close results.
What it does: Compares a deposited vocabulary-learning study's published statistics with an independent re-run, for readers judging reproducibility of word-length effects.
Demo badge: published
Headline result: Fifteen of nineteen published comparisons reproduced with matching direction and significance; four were labelled close.
Start here: Review the study scope—88 students, 16 words—then follow the three models through to the comparison table.
The data:
The dataset comes from a published study by Dr Yen Na Yum (EdUHK) and co-author Wang (2025), deposited publicly on OSF. It covers 88 Chinese-English university students learning 16 technical English words under different reading conditions — the data and original analysis scripts are both openly available. Because the dataset is fixed and fully public, results are stable; the numbers you see today will be the same tomorrow.
What you'd see, step by step:
Try or inspect it:
Inspect the word-length comparison table and trace an effect to the published OSF data. Contact the team for a run on other material under an agreed data-handling arrangement.
Limitations:
The study is relatively small (88 students, 16 words), so effect sizes should be treated as indicative rather than definitive. The replication uses a different but comparable statistical method to the original (the two approaches measure subtly different things), which is why some coefficient numbers differ even when the conclusions agree — sign and significance, not exact magnitude, are the honest replication target. Four of the 19 comparisons are flagged 'close' rather than 'reproduced' because one significance threshold shifted; the headline findings are unaffected.
Re-run two deposited experiments and inspect whether the reported statistical pattern survives modest resampling variation.
What it does: Re-runs two deposited emoji-and-creativity experiments and compares reported mediation results, for readers checking statistical reproducibility.
Demo badge: published
Headline result: Mediation results showed emojis reducing objectification feelings and lifting creativity, with all flagged outcomes matching the published pattern.
Start here: Compare the two experimental samples, then follow exclusions through group tests and mediation to the summary panel.
The data:
Two small experimental datasets (roughly 150–160 participants each) from Prof Sara Kim's published PLOS ONE study, downloaded directly from the open-science repository OSF. The data are real experimental responses, not synthetic, and are frozen at the published version — so results are stable and will look the same every time the demo runs.
What you'd see, step by step:
Try or inspect it:
Re-run the mediation step with a different random seed, then compare the confidence interval and conclusion with the bundled result.
Limitations:
Both studies are small (under 160 participants each) and used student samples in a controlled online setting, so effect sizes may not transfer directly to real workplaces. The confidence intervals reported are ranges, not single definitive numbers — a slightly different resampling method shifts the interval endpoints a little, though the direction and significance of every finding remain the same. This demo shows statistical reproducibility only; it cannot speak to whether the original experiment would replicate in a fresh sample.
Re-run the archived interpreting-error analyses and compare the direction and relative importance of effects in a small exploratory sample.
What it does: Reproduces regression analyses from a consecutive-interpreting error study with 53 students, for readers checking whether effect directions match the paper.
Demo badge: published
Headline result: All fifteen published findings from the paper's main results table reproduced, matching direction across error types and predictors.
Start here: Review the fifty-three interpreters and four error types, then follow collinearity checks through regressions to the comparison table.
The data:
The dataset comes from the paper's own openly shared archive (OSF): error-rate recordings from 53 student interpreters, scored across four error types (conceptual, lexical, syntactic, phonological) alongside each student's language proficiency, working-memory span, and anxiety score. The data are real but small — 53 participants — so findings should be read as exploratory rather than definitive. Because the dataset is a fixed, archived file (not a live feed), results are stable and fully reproducible each time the demo runs.
What you'd see, step by step:
Try or inspect it:
Inspect the regression results table and point to an estimate whose uncertainty you want explained. For another dataset, contact the team first under an agreed data-handling arrangement.
Limitations:
The sample is small (53 students), so individual regression estimates carry meaningful uncertainty — treat the direction of effects as the headline finding, not the precise numbers. The relative-importance scores (e.g. proficiency contributes roughly 18% of explained variance in total errors) are summaries of a pattern, not exact quantities to over-interpret. The replication re-expresses the authors' original R analysis in Python and matches it closely, but cannot go beyond what the original study's design allows.
Re-run the deposited main comparisons across Cantonese, English and Mandarin while flagging the missing items that block the reliability analysis.
What it does: Re-runs language-attitude comparisons across Cantonese, English, and Mandarin ratings from 540 participants, for readers checking main published effects.
Demo badge: published
Headline result: Seven of twelve traits showed clear language differences; English led on intelligence and prestige, Mandarin on employability.
Start here: Review how five hundred forty participants rated one speaker across three language conditions, then open the results table.
The data:
The study asked 540 Hong Kong participants to rate a single recorded speaker — in either Cantonese, English, or Mandarin — on personality and credibility traits using 7-point scales. The dataset is publicly available on OSF (open-access, with a PLOS data-quality badge) and is fixed, so results are stable and reproducible to the decimal point. One caveat: the deposited file appears to be missing some of the original survey items, which affects one secondary analysis (reliability scores) but not the main findings.
What you'd see, step by step:
Try or inspect it:
Inspect the displayed results table, choose a language condition, and ask what a group difference means. For a run on other material, contact the team to agree data handling first.
Limitations:
The original study used only 540 participants in one city at one point in time, so findings should not be generalised beyond Hong Kong or this period. The deposited data is missing some survey items, meaning the authors' internal reliability scores (Cronbach's α) cannot be reproduced — a genuine transparency gap in the public archive, though it does not affect the main statistical conclusions. Participant demographics were not fully balanced across language groups, a limitation the original authors themselves acknowledge.
Compare nine deposited-data tests across four studies, including the substantially smaller Hong Kong effect and software-related numerical differences.
What it does: Re-runs four COVID-era self-concept studies from one deposited paper, for readers comparing reproduced correlations across US and regional samples.
Demo badge: published
Headline result: Role disruption predicted feeling inauthentic more strongly among people who valued those roles highly; nine key statistics reproduced.
Start here: Compare the four samples, then follow filters and regressions to the side-by-side results summary.
The data:
The demo uses the publicly archived datasets from Prof Amy N. Dalton's PLOS ONE paper (Liu, Dalton & Lee, 2021), deposited openly on OSF. The data are real survey responses collected during COVID-19 — two US samples, one US experiment, and one Hong Kong sample — totalling around 1,200 participants across four studies. Because these are fixed, archived files (not live feeds), results are stable and fully reproducible each time the demo runs.
What you'd see, step by step:
Try or inspect it:
Change the bundled filter—for example, retain participants who failed one attention check—then compare the updated sample sizes and results.
Limitations:
All data come from one published study, so this demo shows reproducibility of existing findings, not new discovery. The HK sample effect (r ≈ .21) is noticeably smaller than the US sample (r ≈ .59), which the paper itself acknowledges — the demo does not paper over this. Minor numerical differences between the original SPSS output and this pipeline's output are expected and normal (different software, slightly different standard-error conventions); the substantive conclusions are unchanged across all nine tests. Results should be read as a range of plausible values, not single definitive numbers.
Explore how a synthetic proposal summary can be mapped to published grant criteria for structured pre-submission review.
What it does: Maps a synthetic grant-style summary against published funding criteria, for applicants exploring structured pre-submission self-review.
Demo badge: synthetic
What this demo produces: The demo produces a colour-coded criteria matrix with evidence quotes and flagged gaps for the bundled synthetic summary.
Start here: Review the labelled synthetic proposal summary, then inspect the criteria matrix and revision suggestions.
The data:
The demo uses a clearly labelled synthetic summary of a Hong Kong adolescent sleep and well-being study, written in the style of a typical RGC GRF application. It is not a real proposal. Because the underlying model and live data can vary, results are shown as ranges, not single fixed numbers.
What you'd see, step by step:
Try or inspect it:
Inspect the criteria-alignment table and point to a score or evidence gap you want explained. For another application, contact the team first under an agreed data-handling arrangement.
Limitations:
The demo uses a synthetic summary, not a real application, so scores reflect only what that summary evidences. The scoring model is a proxy reviewer and can vary slightly between runs; the keyword baseline is simple. Always check the current RGC guidance, as wording may change.
Re-run the deposited word norms and compare crowdsourced with AI-generated estimates, while marking analyses that require unavailable external data.
What it does: Compares crowd and machine print age-of-acquisition estimates for about 11,000 English words, for readers evaluating norm coverage and agreement.
Demo badge: published
Headline result: Average learned-in-print ages of about 8.2 for reading and 8.9 for writing match the published paper exactly.
Start here: Open the coverage summary, then compare crowd averages with the deposited paper values.
The data:
The study uses two public datasets: roughly 11,000 common English words rated by nearly 800,000 crowd workers on Amazon Mechanical Turk (asking 'at what age did you first learn this word in print?'), plus a matching set of zero-shot estimates produced by GPT-4o for the same words. Both datasets are openly deposited on OSF by the authors and are stable snapshots, so the numbers you see in the demo will be consistent across runs.
What you'd see, step by step:
Try or inspect it:
Pick any word from the list on screen and ask 'what age did GPT-4o assign this word versus the crowd?' — the demo will pull up both values instantly.
Limitations:
The data covers English words only and the crowd sample skews toward US adults, so norms may not transfer directly to other languages or populations. The AI-versus-human correlations are moderate (roughly 0.67–0.70), meaning AI estimates are a useful but imperfect proxy for human judgements. Some regression analyses in the original paper require an additional external dataset not included in the deposit, so those tables are noted as a follow-up rather than shown live. All figures reported are ranges across slightly different word-overlap subsets, not a single definitive number.
The pipeline re-runs the core statistical analyses from a 2022 PLOS ONE study on child-care worker burnout — and checks how closely the numbers match what was published.
What it does: Re-runs core analyses from a child-care burnout study on its public longitudinal file, for readers comparing reproduced estimates with published values.
Demo badge: published
Headline result: Behavioural-problem trajectories and burnout regressions landed close to published values, with one phase-two effect flagged as divergent.
Start here: Check the longitudinal structure—381 children, 76 workers—then follow descriptives through growth curves to burnout models.
The data:
The dataset is the publicly released SPSS file attached to the original paper (open licence, CC-BY 4.0): 381 children living in Hong Kong residential care homes, assessed repeatedly over two years by 76 care workers. Because this is a fixed, archived dataset — not a live feed — the numbers are stable and will look the same every time the demo runs.
What you'd see, step by step:
Try or inspect it:
Inspect the replication results table and point to an estimate you want compared with the published result. For another dataset, contact the team first under an agreed data-handling arrangement.
Limitations:
The dataset is small (76 workers) and covers only one sector in Hong Kong, so results should not be generalised broadly. The replication used a standard multilevel approach rather than the original software (Mplus), which means slope values look numerically different until a simple scaling adjustment is applied — the direction and significance still match. One phase-2 effect appeared marginally significant in the replication but not in the paper; this is flagged honestly rather than explained away. Effect-size ranges, not single 'best' numbers, are the appropriate way to read the burnout results (roughly 25–33% of variance explained).
Re-run the deposited clinical-survey analyses and inspect where an open-source approximation differs from the original invariance test.
What it does: Re-runs stroke quality-of-life analyses from a published clinical survey file using open tools, flagging where invariance tests differ from the original.
Demo badge: published
Headline result: Descriptive summaries, reliability at omega 0.85, and two-month stability at r equals 0.67 matched the published paper.
Start here: Review baseline and two-month sample sizes, then follow descriptives through factor structure to the validity table.
The data:
The data come from the published supplementary file of Fong, Lo & Ho (2023, Scientific Reports), released under an open licence. It contains questionnaire responses from 184 Hong Kong stroke survivors at baseline and 148 of the same participants two months later, covering quality of life, physical health, mood, hope, self-esteem, and disability. Because the dataset is fixed and public, results are stable — but small real-world samples like this always carry some uncertainty, so the pipeline honestly reports ranges rather than single headline numbers.
What you'd see, step by step:
Try or inspect it:
Inspect the open-tool results table and point to a result you want compared with the published analysis. For another dataset, contact the team first under an agreed data-handling arrangement.
Limitations:
The pipeline uses a freely available Python tool rather than the specialised software (Mplus) used in the original study. This means one technical step — testing whether the scale measures time points on an identical numerical scale — can only be approximated here, not exactly replicated. Fit index numbers therefore differ slightly from the paper's (the pattern is the same), and that gap is flagged openly throughout the demo. The sample is also small (184 people), so individual numbers should be read as indicative ranges.
Inspect whether an exploratory pipeline recovers the planted three-factor structure, then review its draft Methods and Results wording.
What it does: Runs exploratory factor analysis on synthetic survey responses with a known three-factor structure, for researchers testing questionnaire workflows.
Demo badge: synthetic
Headline result: Factor analysis assigned every item to its planted group with one hundred percent match and subscale reliabilities above 0.70.
Start here: Review the planted three-factor design, then follow reliability scores through factor assignment to the draft paragraphs.
The data:
The demo uses a fully synthetic dataset: 240 simulated respondents, 12 questions, answered on a 1–5 scale, with a known three-group structure built in from the start. Because the correct answer is known in advance, we can prove the pipeline finds it before anyone trusts it with real data. A real deployment would use your own survey responses instead.
What you'd see, step by step:
Try or inspect it:
Select one of the bundled scale-label sets and watch the draft update with those fixture names.
Limitations:
The data here is synthetic and tidy — real surveys bring missing responses, scale bunching at one end, and cross-cultural wording effects that can lower these numbers. The factor analysis shown is exploratory only; a journal submission would also need a confirmatory stage. The AI-generated paragraphs are a first draft: the interpretation, emphasis, and final wording must come from you as the author.
Re-run the deposited classification scores across 20 splits, while keeping the unavailable raw-text processing stage explicit.
What it does: Re-runs entropy-based classification on one thousand deposited Chinese texts, for readers checking whether translated versus original patterns reproduce.
Demo badge: published
Headline result: Support-vector and linear classifiers reached roughly eighty-eight to ninety percent accuracy across twenty random splits.
Start here: Inspect the seven deposited entropy scores per text, then follow training splits through accuracy tables to stability panels.
The data:
The dataset is 1,000 short Chinese texts — 500 originally written in Chinese and 500 translated into Chinese from English — drawn from two established public corpora and deposited openly on OSF under a Creative Commons licence. Each text is described by seven numerical scores measuring the statistical 'randomness' of its word and character patterns (Shannon entropy). The dataset is fixed and offline, so results are stable; reported figures are ranges across 20 runs rather than a single lucky number.
What you'd see, step by step:
Try or inspect it:
Inspect the classification table and compare rows for original and translated Chinese texts. For a run on other texts, contact the team first under an agreed data-handling arrangement.
Limitations:
The study's original text-processing step (segmenting raw texts and computing entropy scores) cannot be re-checked because the public deposit contains only the final scores, not the raw texts or processing code; this demo reproduces the classification stage only. The original paper did not record its random split, so an exact number-for-number match is impossible — the right comparison is whether the pattern of results holds, and it does. All accuracy figures are ranges across 20 runs, not single best-case numbers. The dataset is modest in size (1,000 texts) and covers two specific genres, so conclusions may not generalise to all Chinese text types.
Re-run the deposited pretest correlations and regressions for 35 children, while clearly separating statistical agreement from generalisability or diagnosis.
What it does: Reproduces pretest-only correlations and regressions from a dyslexia reading study with 35 children, separating statistical match from generalisability claims.
Demo badge: published
Headline result: All ten published pretest statistics reproduced exactly from the deposited public file.
Start here: Note the evidential boundary—35 children, pretest only—then compare correlations and regressions with published values.
The data:
The dataset is 35 Hong Kong Chinese children with developmental dyslexia, drawn from the pretest phase of a 2021 PLOS ONE study by Prof Yetta Kwailing Wong (CUHK Educational Psychology) and colleagues. It is real, publicly deposited on OSF, and fixed — it does not change between runs, so results are stable and fully reproducible from the deposit. Post-training outcomes were not included in the public file, so only the pretest correlations and regressions can be checked here.
What you'd see, step by step:
Try or inspect it:
Remove one bundled control, such as IQ, then compare how the estimated contribution of character fluency changes.
Limitations:
Only 35 children, one school context, pretest data only — effect sizes this large are unlikely to replicate in a broader or more diverse sample. The pipeline reproduces the paper's own numbers perfectly, but that tells us the analysis was done correctly, not that the findings will generalise. Post-training outcomes and the full longitudinal data are not in the public deposit and cannot be checked here.
Reanalyse three open behavioural datasets with a population-averaged model and compare directions and significance patterns with the paper's subject-specific analysis.
What it does: Reanalyses three deposited reaction-time experiments with a population-averaged model, for readers comparing direction patterns with the published mixed-model results.
Demo badge: published
Headline result: The population-averaged reanalysis agreed with the paper on headline direction and significance patterns across all three experiments.
Start here: Open the archived experiment files, then follow cleaning counts through models to the side-by-side pattern summary.
The data:
The data come from three real behavioural experiments (visual search reaction-time tasks) published in 2025 by Prof Shelley Xiuli Tong's lab at HKU, with all files openly deposited by the authors on OSF. The pipeline works from those original files directly, so the numbers shown are grounded in a real published study, not synthetic examples. Because the pipeline re-estimates the models each time it runs, reported results are given as ranges rather than single fixed figures — minor variation is expected and normal.
What you'd see, step by step:
Try or inspect it:
Inspect the results table, choose one experiment, and compare the reanalysis pattern with the published result. For another dataset, contact the team to agree data handling first.
Limitations:
The pipeline uses a different statistical method (population-averaged GEE) from the one in the paper (subject-specific mixed model), so exact numbers will not match — direction and significance pattern are the right things to compare. Data-trimming percentages come close but not always exactly, reflecting small differences in how ambiguous cases are handled. With only a few dozen participants per experiment, effect sizes are modest and some secondary results sit near the significance boundary, so treat any single p-value with caution.
Re-run selected distribution and growth checks on the archived Korean hip-hop dataset and compare the resulting pattern with the published analysis.
What it does: Re-runs distribution and growth-model checks on a deposited hip-hop collaboration network, for readers comparing patterns with the published analysis.
Demo badge: published
Headline result: A power-law collaboration curve was rejected while an alternative log-normal curve fit better, matching the published comparison.
Start here: Inspect the archived artist network subset, then follow curve fitting through bootstrap tests to the model-comparison table.
The data:
The dataset comes directly from the authors' own public archive (OSF): 3,694 Korean hip-hop artists, with each artist's number of featuring collaborations and a quality score built from likes and prior fame. The data is cleaned and anonymised by the original authors and is freely reusable. Because it is a fixed, archived snapshot, results are stable across re-runs — no live-data drift.
What you'd see, step by step:
Try or inspect it:
Inspect the model-comparison table and compare the power-law and log-normal results; the coauthorship analysis is not included. For another dataset, contact the team to agree data handling first.
Limitations:
Bootstrap p-values shift slightly each run by design; the reliable target is the direction (power-law rejected, log-normal plausible), not an exact decimal. The 'S-shaped' growth pattern described in the paper belongs to the academic coauthorship data, not the hip-hop data — that component was not re-run here. The simulation experiments from the paper (synthetic network models) were also not reproduced in this demo. The dataset is relatively small for network science and covers one music scene, so findings may not generalise beyond Korean hip-hop.
Compare selected group-level patterns from an open child-language dataset using an approximate model; this is not a diagnostic or exact software replication.
What it does: Compares group-level sentence-comprehension patterns from an open child-language dataset using an approximate model, not individual diagnosis.
Demo badge: published
Headline result: Subject relatives were easier than object relatives, and children with language disorder scored lower than typical peers—all matching published directions.
Start here: Review the open dataset for sixty-six children, then compare reproduced group patterns with the paper's headline results.
The data:
The dataset is 66 Cantonese-speaking children (children with Developmental Language Disorder, age-matched typical peers, and younger typical peers) tested on sentence-comprehension tasks; it is the authors' own real data, openly shared on the OSF repository under a Creative Commons licence. Because the pipeline pulls directly from that public archive, results are stable — the source files do not change. Exact numbers may land in a small range across runs due to differences in statistical software, so the demo reports ranges rather than a single pinpoint figure.
What you'd see, step by step:
Try or inspect it:
Select a headline group-level result and request a plain-language methodological gloss. The response must not classify, diagnose, or recommend action for an individual child.
Limitations:
The dataset is small (66 children, one study) and cannot support broad generalisations on its own. The replication uses a different statistical method from the original (a Python approximation rather than the exact R model), so coefficient sizes shift slightly — direction and significance are what reproduce, not every decimal place. A word-for-word software replication would require running the authors' own R script, which is included in the demo package but not shown live.
Vary cleaning and analysis decisions in simultaneous EEG and eye-tracking data, then inspect how an underpowered natural-reading result changes.
What it does: Processes simultaneous EEG and eye-tracking from natural reading in twelve adults, for researchers exploring underpowered word-frequency effects honestly.
Demo badge: published
Headline result: The pre-specified word-frequency N400 contrast was null at p equals 0.797 with twelve participants, as reported with power caveats.
Start here: Follow loading and cleaning with fixation–EEG alignment visible, then inspect the N400 contrast and validation quality checks.
The data:
ZuCo (Zürich Cognitive Language Processing Corpus) 1.0 — the natural-reading task: 12 healthy adults read English sentences while EEG and eye-tracking were recorded simultaneously. The public research data was mirrored and verified from the original distribution; this demo processes 36,562 word-level fixation-locked FRP epochs through a full MNE-Python pipeline (band-pass filter, ICA, fixation-to-EEG synchronisation at 2 ms/sample precision).
What you'd see, step by step:
Try or inspect it:
Open the Explore assumptions tab, then change the N400 window, subject-exclusion or sync thresholds and review the recomputed statistics. Non-default runs are marked EXPLORATORY.
Evidence:
ZuCo 1.0 natural-reading, 12 participants, 36,562 word-level fixation-locked FRP epochs (EEG + eye-tracking synchronized at 2 ms/sample). Headline frequency effect is null — p=0.797, d=−0.076, n=12 — reported honestly with power caveats (≈9% at the pre-specified bound). The Explore assumptions tab lets visitors vary the N400 window, subject-exclusion and sync thresholds with live client-side recomputation; non-default runs carry an EXPLORATORY banner and underpowered analyses are blocked. validated 2026-09-01.
Limitations:
With 12 participants the study is underpowered for confirmatory claims — all inferential results are descriptive (power ≈ 9% at the pre-specified bound). The frequency split discards graded frequency information, sentence context was not controlled, and the next-fixation response overlaps the analysis window (fixation-overlap contamination risk). One subject's run was recovered via alternate runs; one run is absent from the mirror and disclosed. The cross-modal correlation is model-dependent and cannot be confirmed at this sample size.
Recalculate selected evaluation metrics from shared Shenzhen image ratings and precomputed AI predictions, while flagging an unresolved baseline-version discrepancy.
What it does: Recalculates evaluation metrics from shared street-image ratings and stored predictions, for readers comparing model scores with survey restorative judgments.
Demo badge: published
Headline result: Vision-method scores reached R-squared 0.76 with built-in urban knowledge versus 0.37 without, matching reported values from the shared files.
Start here: Review the five-hundred-sixty-six-image dataset and stored predictions, then compare human ratings with recalculated model scores.
The data:
The study uses 566 street-view photographs of Shenzhen collected via a mapping service, each rated by survey respondents for how psychologically restorative the space feels. The images and ratings are publicly available on GitHub alongside the original analysis code. Because the publicly shared files appear to differ slightly from the exact version used to produce the published results, some numbers come out as ranges rather than single figures — that gap is itself part of what the demo illustrates.
What you'd see, step by step:
Try or inspect it:
Inspect the audit table and compare survey ratings with the image-model predictions for a bundled street image. For another dataset, contact the team to agree data handling first.
Limitations:
The dataset is small (566 images, one Chinese city, one season) so results should not be generalised to other cities or climates. The AI vision predictions themselves were made by a commercial model and are taken from the published files — the demo verifies the scoring of those predictions, not the full end-to-end pipeline. Most importantly, the paper's headline claim of a 0.535 improvement over the baseline cannot be independently confirmed yet: the baseline score differs between the published notebook and the shared data files, likely due to a data-version change. The improvement is probably real, but the exact magnitude is unverified.
Independently recode 77 public news articles and inspect where automated labels agree—or disagree—with a transparent keyword check.
What it does: Independently codes seventy-seven news articles on inclusive education framing, for media researchers comparing automated labels with published category splits.
Demo badge: published
Headline result: Both codings placed efforts first at roughly sixty-five percent published versus sixty-one percent in this pipeline's independent labels.
Start here: Compare counting units—concordance instances versus sentences—then inspect flagged sentences and category splits.
The data:
The corpus is 77 official English-language news articles on inclusive education in China, drawn from a public archive deposited by the paper's authors on OSF. The texts are real published news items, not synthetic, and the deposit is stable — the same file is retrieved every run. Because the authors' original line-by-line coding sheets were not deposited, the pipeline builds its own independent coding for comparison rather than replaying the authors' exact work.
What you'd see, step by step:
Try or inspect it:
Select a bundled news-sentence fixture and inspect its proposed category and reasoning.
Limitations:
The authors' original sentence-by-sentence coding was never made public, so this demo cannot reproduce their exact numbers — it produces an independent re-coding for comparison only. The two codings use different counting units (the paper counts 520 concordance instances; the pipeline counts roughly 1,000 sentences), which partly explains differences in the minority categories. The automated coder agrees only moderately with a simple keyword check, meaning borderline sentences — especially praise-heavy policy lines — could reasonably be labelled 'efforts' or 'consensus' by different coders, human or machine. Finally, the corpus covers only official-channel English reporting, so findings describe how authorities project inclusive education abroad, not the full range of media coverage.
See how an AI applies a fixed category codebook to 30 public complaint narratives and compare its classifications with the dataset's existing labels.
What it does: Applies a fixed five-category codebook to thirty public complaint narratives, for researchers comparing AI labels with existing regulator categories.
Demo badge: real
Headline result: Twenty-seven of thirty texts matched the source labels, yielding about ninety percent agreement and kappa near 0.73.
Start here: Review the bundled public narratives and codebook, then compare each AI assignment with the source label.
The data:
Thirty real consumer-complaint narratives from a US financial regulator's public database, standing in for open-ended survey or interview responses. The regulator's own category labels serve as the human-coded ground truth. Because this is a fixed public dataset the results are stable, but if you substituted live or updated text the model's agreement scores would shift — treat any figures as a range, not a fixed number.
What you'd see, step by step:
Try or inspect it:
Select an included public complaint, inspect the assigned category, and compare it with the source label. Do not paste confidential interview, survey, or participant text.
Limitations:
The dataset is small (30 items) and comes from one specific domain, so the kappa figure (roughly 0.70–0.80 across runs) is illustrative rather than definitive. Categories with only one or two examples — such as 'credit card' or 'debt collection' here — can show wide swings in agreement by chance. A real study would need a larger, domain-matched sample and a codebook refined through the disagreement-review step shown in the demo.
Explore how a real single-subject EEG recording is cleaned, averaged, and visualised—and where a tutorial workflow stops short of study-level evidence.
What it does: Traces a single-subject EEG recording from raw traces to event-related potentials and scalp maps, for students learning standard preprocessing steps.
Demo badge: real
What this demo produces: The tutorial pipeline cleaned roughly ninety percent of trials and produced separable auditory versus visual averaged waveforms.
Start here: Inspect the raw sixty-channel recording, then follow filtering, epoching, and averaging to the scalp maps.
The data:
The demo uses a standard open-access teaching dataset: a 60-channel EEG recording of a person responding to auditory and visual cues, roughly two and a half minutes long. It is a real (not synthetic) recording widely used in neuroscience training worldwide, so results are stable and repeatable across runs. Because this is a single-subject tutorial file rather than a full study dataset, the numbers shown are illustrative benchmarks, not publishable group statistics.
What you'd see, step by step:
Try or inspect it:
Point to a brain-map or signal panel and ask what it shows for a reading or language study. Contact the team for another run under an agreed data-handling arrangement.
Limitations:
This is a single participant from a tutorial dataset, so no group statistics or individual differences are shown — a real study would add per-participant tuning and mixed-effects models. The pipeline settings (filters, epoch length, artefact thresholds) are reasonable defaults, not optimised for a specific research question; a live project would involve a pre-registered analysis plan. Results should be read as a proof-of-concept demonstration of the workflow, not as findings.
Test how a Cantonese-language prototype routes 15 scripted messages through non-diagnostic distress flags and crisis-handoff rules.
What it does: Routes fifteen scripted Cantonese messages through distress flags and crisis handoff rules, for teams reviewing triage safety—not live counselling.
Demo badge: synthetic
Headline result: All fifteen distress levels were identified correctly and every crisis message triggered a hotline number and human-handoff recommendation.
Start here: Select scripted everyday, distress, and crisis examples only—do not enter personal experiences or identifying details.
The data:
Fifteen scripted Cantonese student messages, split evenly across everyday chat, moderate distress, and crisis statements (including self-harm and suicidal ideation). The messages are purpose-written for this demo, not drawn from real student records — this is a proof-of-concept mechanism test, not a clinical study. Because the scripts are fixed, results are stable across runs; a live deployment with real users would behave differently and would need proper ethics-approved evaluation before any conclusions could be drawn.
What you'd see, step by step:
Try or inspect it:
Use the supplied selector to compare everyday, moderate-distress, and crisis-handoff scripts. Do not enter personal experiences, identifying details, or requests for crisis support.
Limitations:
All 15 messages were written by the research team, so the chatbot has never been tested on real students; strong results here confirm the design works in principle, not in practice. The automated judge is itself an AI, so empathy scores are indicative, not clinically validated — the one hard check is a simple text-search confirming the real hotline number appears in every crisis reply. This is a triage-and-handoff tool only: it does not diagnose, treat, or counsel, and any live deployment would need ethics approval, clinician review of the scripts, and verification of local crisis resources.
Explore sentiment calibration, theme extraction and summary drafting on a fixed English review corpus, with human review required for transfer.
What it does: Calibrates sentiment, extracts themes, and drafts summaries on a fixed English review corpus, for analysts learning text-mining with human-check labels.
Demo badge: real
Headline result: Sentiment labels matched human judgments at ninety-six percent accuracy on the one-hundred-twenty-review calibration subset.
Start here: Start with the labelled calibration subset within the two-thousand-review corpus, then inspect sentiment accuracy before theme extraction.
The data:
The demo uses 2,000 real IMDB consumer reviews drawn from a publicly available dataset, with a 120-review sub-sample that carries human-verified sentiment labels used to check the AI's accuracy. Because this is a fixed public dataset the results are stable for today's walkthrough; if the same pipeline were pointed at live social-media or regulatory feeds, scores would shift over time and should be treated as ranges rather than single fixed numbers.
What you'd see, step by step:
Try or inspect it:
Select a bundled review or policy-text fixture and inspect its sentiment label and nearest theme.
Limitations:
The accuracy figure (96%) comes from English-language reviews; Chinese or Cantonese text needs additional language-specific setup before the same reliability applies. Topic names are generated from the most frequent words in each cluster — a human researcher should review and rename them for any publication. The pipeline is validated on a 120-review sample; a full deployment on your own corpus would require a fresh calibration run on that material.
Trace a synthetic reaction-time experiment from trial-level checks to statistics and an author-reviewed draft Results paragraph.
What it does: Traces a simulated reaction-time experiment from trial data through exclusion rules to a draft Results paragraph, for cognitive researchers testing analysis flow.
Demo badge: synthetic
Headline result: Related-prime responses averaged forty-five milliseconds faster than unrelated primes in the built-in synthetic dataset.
Start here: Inspect the raw trial-by-trial response times, then follow exclusion rules through group averages to significance output.
The data:
The demo uses a simulated word-response experiment: 32 participants each completing 80 trials, with a known 45 ms speed difference built in so the pipeline can be checked against a correct answer. The data are synthetic (not from a real study), designed to mimic standard cognitive-psychology experiments. Because the data are fixed for this demo, results are stable — a live pipeline on real collected data would produce ranges rather than single numbers.
What you'd see, step by step:
Try or inspect it:
Inspect the bundled trial results and exclusion summary, then point to a rule you want explained. To run your own study, contact the team first under an agreed data-handling arrangement.
Limitations:
The data are simulated, so the pipeline has not yet faced the messier trial structures of a real study — though the steps carry over directly. Only a simple two-condition comparison is shown here; real multi-experiment papers typically need more complex models, which are an extension not yet in this demo. The AI-drafted paragraph is a writing aid and starting point, not a submission-ready text.
Compare AI judgements across ten controlled translation pairs, including preference, error detection and less-reliable error-type labels.
What it does: Scores ten Chinese–English translation pairs with planted errors, for teams testing whether an AI judge prefers the polished version and spots flaws.
Demo badge: synthetic
Headline result: The judge preferred the good translation in all ten cases and spotted the planted flaw in nine of ten.
Start here: Review the ten sentence pairs, then inspect adequacy and fluency scores beside each flagged error type.
The data:
Ten Chinese–English sentence pairs were hand-crafted for this demo: each pair contains one polished translation and one with a single planted error (mistranslation, omission, unnecessary addition, awkward phrasing, or wrong register). The data is small and purpose-built to test the mechanism, not to benchmark real-world performance. Because this is a fixed test set rather than live text, scores are stable across runs.
What you'd see, step by step:
Try or inspect it:
Choose one of the bundled held-out translation pairs and inspect the score, preference and flagged concern.
Limitations:
This is a 10-sentence proof-of-concept with errors that were deliberately planted, so the strong numbers (100% preference accuracy, 90% error detection) reflect a controlled setting, not a live translation workflow. Error-type labelling is less reliable — the AI correctly named the error category about two-thirds of the time, and occasionally confused related error types. AI scoring correlates with human judgment but should be treated as a first-pass triage tool, not a replacement for human assessment in high-stakes contexts. A real study would use 100+ sentences evaluated by at least two human raters alongside the AI.
Use bundled synthetic fixations to inspect reading-measure calculations and a draft paragraph that still requires author review and study-level modelling.
What it does: Computes standard reading measures on synthetic Chinese fixation data with known frequency effects, for readers validating measure calculations.
Demo badge: synthetic
Headline result: Low-frequency words took longer to read on every computed measure, matching the planted frequency difference.
Start here: Open the bundled fixation records, then follow measure calculation through to the known-effect comparison.
The data:
The demo uses simulated Chinese sentence-reading data (30 readers, 40 sentences each) with known word-frequency differences built in, so the pipeline can be checked against a correct answer. This is synthetic validation data, not a real corpus; real deployment would use your lab's own recordings or public datasets such as ZuCo or GECO. Because the demo data is fixed, the numbers you see today will be stable — but results on live lab data will naturally vary across runs and participant samples.
What you'd see, step by step:
Try or inspect it:
Choose a bundled word-level fixation fixture and re-run the calculations using the supplied column mapping.
Limitations:
The data today is synthetic and small (30 simulated participants), so this validates that the measures are computed correctly, not that the norms match any real population. The pipeline does not yet include mixed-effects models, which are the current statistical standard for trial-level eye-tracking data — the paired t-tests shown are a useful first pass only. The AI-drafted paragraph is a starting point; authors must review it before submission.
Inspect how a synthetic Cantonese fixture becomes a word-level transcript and pitch table—and where natural audio would require correction and additional alignment.
What it does: Transcribes a synthetic Cantonese passage and extracts word-level pitch and timing, for speech researchers testing prosody measurement on clean audio.
Demo badge: synthetic
What this demo produces: One hundred thirty-five aligned units were processed with pitch measurements at about 3.0 units per second in the bundled clip.
Start here: Play the bundled Cantonese clip, then inspect the transcript and per-word pitch table as they appear.
The data:
The demo uses a short synthetic Cantonese passage (45.6 seconds, 135 aligned units) generated by a text-to-speech voice, so results are perfectly reproducible for a live showing. Real interview or corpus recordings can be substituted directly — the pipeline handles them the same way. Because synthetic speech is cleaner than natural conversation, pitch and timing figures on real recordings will vary; treat any numbers as illustrative ranges, not fixed benchmarks.
What you'd see, step by step:
Try or inspect it:
Choose another bundled Chinese sentence fixture and re-run the pipeline; compare its transcript and pitch table with the supplied reference.
Limitations:
The audio here is synthetic (computer-generated speech), so it is unusually clean; natural recordings — with background noise, overlapping speakers, or heavy dialectal variation — will produce messier pitch tracks and occasional transcription errors. Timing is at the word level, not the individual sound level, so fine-grained phoneme work would need an extra alignment step. Pitch figures are raw and would require speaker-specific tuning before appearing in a publication.
Compare rate, transcript accuracy and pitch variation across two controlled recordings without treating the outputs as child norms or validated assessment scores.
What it does: Compares rate, accuracy, and expression scores across two synthetic Chinese readings of the same passage, for educators testing fluency measurement—not norms.
Demo badge: synthetic
Headline result: The fluent recording beat the disfluent recording on every expected rate, accuracy, and rubric measure in the self-check.
Start here: Open the passage text and two audio files, then compare transcripts and side-by-side fluency scores.
The data:
Two audio recordings of the same 83-character Chinese passage were used: one read fluently, one read slowly with four deliberate word substitutions (e.g., 森林 replaced by 公園). Both voices are synthetic — real children's recordings would be the next step. Because the conditions are fixed and controlled, results are stable for this demo, but scores on real classroom audio would vary with recording quality and individual children.
What you'd see, step by step:
Try or inspect it:
Ask the demo to play back one of the two recordings while you watch the live transcript appear word by word.
Limitations:
Both recordings use a synthetic (text-to-speech) child voice, so pitch-variety scores do not meaningfully separate the two conditions — that finding needs real expressive versus flat child readings to be trusted. Accuracy figures are a conservative lower bound because the speech recogniser occasionally 'corrects' substituted words back to the original text; a production system would use forced alignment against the passage to catch every error. This is a two-condition demo of the measurement machinery, not a norming study — it shows the pipeline works, not what typical scores look like across a class or grade level.
Test three human-reviewed administrative workflows on fictional materials: email triage, reference-letter drafting, and meeting-note summarisation.
What it does: Tests human-reviewed email triage, reference drafting, and meeting summaries on fictional academic materials, for staff exploring admin workflows safely.
Demo badge: synthetic
Headline result: All ten email categories matched gold labels, but priority suggestions were correct only half the time.
Start here: Review the ten fictional emails, then compare the suggested categories and actions with your own judgment.
The data:
The demo uses ten hand-written realistic academic emails, one fictional student CV (for Chan Hoi Yan, an MPhil Linguistics candidate), and a transcript of one fictional departmental meeting — all synthetic and modelled on a Hong Kong university context. Because the data is fixed and self-contained, results are stable across runs; a real deployment using your own emails would need two to four weeks of tuning and results would vary.
What you'd see, step by step:
Try or inspect it:
Choose a bundled fictional email, edit its wording without adding personal or confidential information, then compare the suggested category and action with your judgment.
Limitations:
All three tasks run on ten emails, one CV, and one meeting — a proof of concept, not a tested system. Gold labels reflect one person's judgment, so accuracy figures should be read as illustrative ranges rather than firm benchmarks. The agent drafts and flags; it does not send emails or submit forms — integration with Outlook or university portals is a separate engineering project not covered here. Priority scoring was only 50% accurate on this small set, so treat priority suggestions as a prompt to check, not a decision.
Inspect turn counts and talk-share estimates from a clean synthetic fixture, without claiming equivalence to a validated research device.
What it does: Measures turn counts and talk share from a synthetic parent–child Cantonese clip, for researchers testing conversational metrics—not validated field data.
Demo badge: synthetic
Headline result: The pipeline detected six child vocalisations across eleven turns with adult talk share about seventy-three percent in the bundled clip.
Start here: Play the bundled forty-eight-second clip, then review turn labels and talk-share estimates against the scrollable transcript.
The data:
A short Cantonese parent–child play session (roughly 48 seconds, 12 scripted turns) was synthesised for the demo by pitch-shifting an adult voice to approximate a child's voice — a standard acoustic technique. The audio is fully synthetic and clean, not drawn from real fieldwork. Because the demo uses a fixed recording, results are stable; a live deployment with real noisy recordings would produce ranges rather than exact counts.
What you'd see, step by step:
Try or inspect it:
Select one of the bundled short audio fixtures and compare how turn counts and talk-share estimates change.
Limitations:
The audio is synthetic and clean — no overlapping speech, no background noise. Real nursery or home recordings are much messier and would require additional tuning before the counts are reliable. Speaker labels here were verified against a script written by the demo author, not independent human coders. Treat the headline numbers as a proof of concept, not a validated research instrument.
A non-diagnostic mechanism test shows which planted pronunciation differences the pipeline detects—and which the recogniser obscures.
What it does: Tests whether a Cantonese speech pipeline spots planted child pronunciation errors, for clinicians and researchers evaluating mechanism—not diagnosis.
Demo badge: synthetic
Headline result: Seven of ten disorder patterns were correctly identified and eight of ten planted errors were confirmed present in transcripts.
Start here: Listen to the adult and child voice passages, then inspect whether deliberate mispronunciations survive transcription.
The data:
Ten target words were recorded using synthetic Cantonese child and adult voices, each word embedded in a standard picture-naming phrase used in real clinical assessments. The voices are computer-generated, not real children — this is a mechanism test, not a clinical dataset. Because the errors and targets are built in by design, we know the exact right answer for every item, which makes the accuracy figures fully verifiable.
What you'd see, step by step:
Try or inspect it:
Point to a word in the results table and ask why it was marked unclear; use the displayed per-item evidence to check the explanation.
Limitations:
All voices are synthetic, so the demo proves the pipeline works in principle — it does not yet measure how well it handles real children's voices, which are less precise and more variable. Ten items is a pilot; a clinically meaningful evaluation would need 100 or more items checked against human speech therapist judgements. The speech-recognition engine occasionally 'corrects' a mispronunciation back to the standard word before the disorder detector even sees it, which is a known technical hurdle this demo is designed to expose, not hide.
See how a short synthetic lesson moves from audio to an illustrative coded transcript.
What it does: Shows how a short synthetic lesson moves from audio to coded discourse turns, for educators exploring automated classroom transcription—not real recordings.
Demo badge: synthetic
Headline result: Automatic discourse codes agreed with a hand-written gold standard on about two-thirds of teacher and student turns.
Start here: Review the bundled synthetic lesson waveform, then follow transcription through turn labels to the reliability table.
The data:
The demo uses a short synthetic Cantonese maths lesson (under a minute, 10 exchanges between one teacher and one student) recorded in a clean, controlled way so the pipeline always runs the same way live. Because it is synthetic, results are stable for this demo; a real classroom recording would introduce background noise, overlapping voices, and children's accents, so accuracy figures should be treated as a best-case range rather than a fixed number.
What you'd see, step by step:
Try or inspect it:
Inspect the bundled synthetic lesson and compare its coded exchanges. Do not upload real classroom data; any real-data run requires prior contact, consent, and agreed data handling.
Limitations:
The audio is synthetic — two clean, adult voices with no overlap — so accuracy will be lower on real classroom recordings with children's voices, noise, and cross-talk. Discourse codes are illustrative; a real study would develop the codebook together with your research team. Treat the two-thirds code-agreement figure as an upper bound, not a guaranteed result.
Compare two detection methods on a small constructed set and inspect why paired performance does not establish single-text reliability.
What it does: Compares two AI-text detection approaches on twenty constructed human-and-machine passage pairs, for readers assessing mechanism limits—not benchmark claims.
Demo badge: synthetic
Headline result: A style classifier scored sixteen of twenty correct; paired comparison judged all ten topic pairs but only eleven of twenty single passages.
Start here: Inspect the twenty tagged passages, then compare classifier and judge results row by row in the summary table.
The data:
The demo uses 20 short passages: one genuine human excerpt (from Wikipedia) and one AI-generated passage on the same topic, repeated across 10 Hong Kong-relevant topics (e.g. Cantonese, MTR, Hong Kong cuisine). Because the labels are known by construction — we made the AI passages ourselves — no human annotation is needed to check whether the detectors are right. The texts are a small proof-of-concept set; a real study would require hundreds of passages across many writing styles and AI models.
What you'd see, step by step:
Try or inspect it:
Point to any topic in the results table and inspect how the judge reached its decision for that passage. Use only the 20 bundled passages.
Limitations:
This is a 20-text mechanism demo, not a benchmark — results would shift with a larger or more varied dataset. The style classifier (16/20) is a weak-but-real signal; the AI judge is stronger at paired comparison (10/10) but unreliable on single passages (11/20), and it tends to flag encyclopedic human writing as AI. Detection accuracy is also a moving target: newer AI models write less predictably, so these numbers reflect today's AI output style and should be treated as indicative ranges, not fixed figures.
A guided pipeline retrieves one completed thesis and associated files, approximately re-runs one verifiable experiment, and compares selected outputs with the thesis.
What it does: Retrieves one completed thesis and associated experiment files, then re-runs one verifiable listening study for approximate comparison—not full replication.
Demo badge: real
Headline result: One experiment's group accuracy and tone-error patterns were approximately reproduced from the supplied raw files; other chapters lacked machine-readable data.
Start here: Use only the supplied thesis and data files, then compare the recomputed statistics table with thesis-reported values.
The data:
The source material is a real, completed PhD thesis (289 pages) on accented-speech intelligibility, submitted to UNSW in 2013, plus four raw experiment data files from the same Dropbox folder — together covering several thousand listening trials across multiple experiments. The data are real and finished (not synthetic), so results are stable; however, the pipeline uses a simplified statistical approach rather than the full mixed-model analyses in the original thesis, which is why some numbers come out slightly different from the published figures. Any figures quoted should be treated as ranges, not single authoritative values.
What you'd see, step by step:
Try or inspect it:
Use only the supplied files: point to a comparison-table row and ask the panel to explain it. For another project, contact the team first to agree scope, permissions, and data handling.
Limitations:
Only one experiment's statistics were fully verifiable from the raw files provided; other chapters had no machine-readable data attached, so the pipeline cannot check those. The statistical tests used here are simplified (basic t-tests on group means) rather than the mixed-model ANOVAs in the original thesis — this explains small differences in t-values and effect sizes, and means results should be read as approximate reproductions, not exact replications. The eye-tracking session log was parsed for trial structure only; actual gaze data requires specialist export files not included. All findings come from a single, decade-old thesis with modest sample sizes (50–60 participants per study), so nothing here should be generalised.