SimCoachCorpus

A naturalistic dataset with language and trajectories for embodied teaching
Accepted to KDD Datasets & Benchmarks Track 2026
Emily Sumner* Deepak E. Gopinath* Laporsha Dees Patricio Reyes Gomez Xiongyi Cui Andrew Silva Jean Costa Allison Morgan Mariah Schrum Tiffany L. Chen Avinash Balachandran Guy Rosman Toyota Research Institute * Equal contribution — joint first authors
A professional driving coach observes and instructs a participant driving in the racing simulator.
SimCoachCorpus captures how people acquire motor skills in an embodied task through verbal instruction over time, synchronizing race-car simulator trajectories with a professional coach's concurrent and terminal feedback.

Abstract

High-quality curated datasets are essential for training and evaluating AI approaches, but are often lacking in embodied interactive domains where language and physical action are intertwined. In particular, few datasets capture how people acquire motor skills in embodied tasks through verbal instruction over time. To address this gap, we introduce SimCoachCorpus: a unique dataset of race car simulator driving that enables the investigation of rich phenomena during guided and unguided motor skill acquisition. In this dataset, 29 humans were asked to drive in a driving simulator around a race track for approximately ninety minutes. Fifteen participants received one-on-one instruction from a professional performance driving coach, and 14 participants drove without coaching instruction. SimCoachCorpus includes features such as vehicle state and inputs, map (track boundaries and race-line), and cone landmarks. Additionally, these are synchronized with the coach's concurrent verbal feedback and additional terminal feedback at the end of each lap. We also provide high-quality annotations of high-level coaching categories for each concurrent feedback utterance, ratings on students' compliance with coaching advice, and self-reported cognitive load and emotional state of participants (gathered from surveys during the study). The final dataset includes over 20,000 concurrent feedback utterances, over 400 terminal feedback utterances, and over 40 hours of interactive driving data. Our naturalistic interactive dataset can be used to investigate motor learning dynamics, explore linguistic phenomena, and train computational models of teaching and learning. We demonstrate applications of this dataset for in-context learning, imitation learning, and topic modeling.

29 participants15 coached, 14 uncoached
40+ hoursof interactive driving
20k+ utterancesconcurrent coach feedback
400+ feedbackat the end of each lap
100+ surveysself-reported cognitive load, emotion & motivation

Dataset

Each participant drove around a race track in a high-fidelity driving simulator for roughly ninety minutes. Vehicle telemetry is synchronized with the coach's speech, so language and physical action can be studied together over time. The dataset includes:

Vehicle state and inputs — full trajectories, controls, and dynamics.
Map information — track boundaries, the optimal race-line, and cone landmarks.
Concurrent verbal feedback — the coach's utterances during driving, time-aligned to trajectories.
Terminal feedback — the coach's summary feedback at the end of each lap.
High-level coaching annotations — category labels for every concurrent feedback utterance.
Compliance ratings — how well students followed the coaching advice.
Cognitive load & emotional state — self-reported survey responses collected during the study.

Study protocol for the coached and self-practice conditions, showing the sequence of surveys, demonstration lap, baseline laps, repeated driving blocks, and retention laps.
Study protocol. Participants in the coached condition received live feedback from a professional driving coach, while those in the self-practice condition drove without feedback. Both groups completed surveys, a demonstration lap, baseline laps, repeated driving blocks, and retention laps.

Applications

SimCoachCorpus can be used to investigate motor learning dynamics, explore linguistic phenomena, and train computational models of teaching and learning. In the paper, we demonstrate three applications:

In-context learning — using coaching context to inform driving predictions.
Imitation learning — learning driving policies from human trajectories.
Topic modeling — uncovering structure in the coach's feedback language.

Hierarchical clustering dendrogram of topics discovered in the coach's concurrent feedback.
Topic modeling of the coach's concurrent feedback. Hierarchical clustering groups related coaching topics — e.g., braking, throttle, steering, racing line, and cone landmarks — revealing structure in the language of instruction.

BibTeX

@article{sumner2025simcoachcorpus, title = {SimCoachCorpus: A naturalistic dataset with language and trajectories for embodied teaching}, author = {Sumner, Emily and Gopinath, Deepak E. and Dees, Laporsha and Reyes Gomez, Patricio and Cui, Xiongyi and Silva, Andrew and Costa, Jean and Morgan, Allison and Schrum, Mariah and Chen, Tiffany L. and Balachandran, Avinash and Rosman, Guy}, journal = {arXiv preprint arXiv:2509.14548}, note = {To appear at KDD 2026, Datasets and Benchmarks Track}, year = {2025} }