Our mission is to improve teaching and learning at scale by learning from great tutors.
We are a team of researchers, educators, and technologists building the world's largest dataset of authentic tutoring interactions, along with the open-source research infrastructure needed to work with it responsibly. We partner equitably with tutoring providers, school districts, and underserved communities.
🌐 nationaltutoringobservatory.org · 💻 GitHub
Sandpiper is our open-source AI textual-annotation application (MIT licensed). It powers the annotation and de-identification workflows behind our datasets, including the proposer–reviewer pipeline we use to detect and replace personally identifiable information in tutoring transcripts.
Our work on utility-preserving de-identification examines a problem specific to math tutoring: numeric expressions often look like structured identifiers, so naive redaction destroys the instructional content it is meant to protect.
📄 Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset — Zhou, Vanacore, Ahtisham, Lee, Pietrzak, Hedley, Dias, Shaw, Schäfer, Kizilcec (arXiv:2602.16571)
The paper introduces MathEd-PII, a companion benchmark dataset for PII detection in math tutoring dialogue.
Our datasets contain authentic educational interactions, so we release them under gated access: the dataset card and terms are readable by anyone, while downloading requires agreeing to our Data Use Terms and a short review of intended use. This lets us share data that would otherwise stay locked away, while keeping commitments to the students, tutors, and institutions who made it possible.
Datasets that are not currently listed here are in preparation or under review. Please get in touch if you have questions about availability.
Where our data has been de-identified using a hide-in-plain-sight approach, identifiers you see in the text are surrogates — realistic but fabricated substitutes. Dataset cards explain this in detail; please read them before use.
General inquiries and data access questions: sandpiper_admin@cornell.edu
This work is supported by the National Science Foundation (Grant No. 2321499), the Gates Foundation, and the Chan Zuckerberg Initiative. Any opinions, findings, and conclusions are those of the authors and do not necessarily reflect the views of the funders.