Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.
Sam: Getting the right notes out of a recording is not the same as getting a score you can read. A new system called SheetSage2 tries to close that gap. It turns a pop song into a coherent lead sheet with a single model. Its training leans heavily on synthesized music rendered from symbolic files.
Alex: Place it for me before we go further. Whose work is this, and what kind of document is it?
Sam: It's a technical report posted to arXiv, so a preprint, not peer reviewed. The team spans New York University, MBZUAI, ACE Studio and the Hong Kong University of Science and Technology, with Yann LeCun and Gus Xia as corresponding authors. Their headline: one model beats the listed prior systems on 12 of 15 benchmark comparisons, across eight benchmark collections.
Alex: I work next door to music research, not in it. Why should someone like me care?
Sam: A lead sheet is the one-page score a band reads. It holds melody, chord symbols, bar lines, key, and section labels like verse and chorus. Producing one from audio has meant gluing separate tools together. This report offers one system, with public weights, that emits the whole thing consistently. So let's cover why gluing fails, the trick that makes it work, and how far to trust the numbers.
Alex: Start with the gluing. If you have a good beat tracker and a good chord estimator, why not just combine them?
Sam: Because the pieces depend on each other. How you spell a note, as G sharp or A flat, depends on the local key. Note durations only make sense on a shared grid of beats and bar lines. Independent predictions don't guarantee that compatibility. The authors' predecessor, SheetSage1, worked roughly that way, combining beat tracking, key, chord and melody estimation.
Alex: Then the obvious move is one sequence model trained end to end. Why didn't anyone just do that?
Sam: Data. You'd need lots of complete, well-aligned pairs of audio and full lead sheets. Those are scarce. Most datasets label only one thing, beats or chords or melody, and even those labels can have sloppy timing.
Alex: So where does the training data come from?
Sam: The first stage is synthetic. MIDI files, which list notes, timing and instruments symbolically, are plentiful. One collection, the Los Angeles MIDI Dataset, has about 400,000 deduplicated files. But their key, chord and melody labels are often missing or wrong. So the authors label them with symbolic models. Then they render the files to audio with a basic software synthesizer, giving aligned audio and labels at scale.
Alex: How do you label melody in a MIDI file you don't trust?
Sam: They bootstrap. They start with a small, high-precision seed set. For melody, that's files where a track is literally named "melody", about 17,000 of them. They fine-tune a pretrained symbolic model on those, then run it across the collection. That expands melody labels to roughly 186,000 songs, nearly eleven times the seed.
Alex: And the second stage?
Sam: A model they call the Prober. It's really two acoustic models built on a pretrained music audio encoder called MERT. Rather than generating a sequence, they predict probabilities frame by frame: beats, downbeats, tempo, chords, key, sections, melody. Frame-wise output tolerates labels that are off by a few milliseconds, and missing labels can simply be masked out. That lets it learn from messy real recordings and synthetic audio together.
Alex: Frame-wise probabilities still aren't a score. What turns them into one?
Sam: That's the central idea: structured decoding. It's a chain of rule-guided searches, each fixing a decision that constrains the next. Tempo first, then beats, then bars and meter. Then key and sections, which may only change at bar lines. Then chords aligned to beats, and finally melody on subdivisions of that beat grid.
Alex: Give me a case where the order actually matters.
Sam: Their example: a song with two verses that share the same rhythm, but only the second has drums. Local predictions might call the first verse 80 beats per minute and the second 160. The music hasn't changed, but now the two verses get different grids and different note lengths. So the decoder first pools tempo evidence over the whole recording and picks one anchor. Local tempo can still drift, but only relative to that anchor.
Alex: So it's a global prior that stops the half-tempo, double-tempo flip-flopping beat trackers are known for.
Sam: Yes. And melody gets a similar fix for octave errors. Instead of scoring each pitch alone, the model scores a pitch together with the interval from the previous note. Say C, E, G gets transcribed with the C an octave too high. That shape is penalized because the jump to E now looks wrong.
Alex: And the third stage, where the single model comes in?
Sam: Distillation, meaning a student learns from a teacher's outputs. They run the Prober and its decoders over about 441,000 recordings, mostly freely licensed and synthetic music. The decoded event sequences become targets for one autoregressive model, which writes tokens for time, meter, key, chord, and melody in one stream. At inference it needs none of those task-specific searches. A deterministic builder converts the stream into notation or a MIDI file.
Alex: Let's get to evidence. Twelve of fifteen is a scoreboard. Which result matters most in practice?
Sam: Melody is the clearest jump. On the RWC pop collection, vocal melody F1 goes from about 63 for SheetSage1 to about 83 for the new model. Downbeats and chords also improve over specialist systems. It doesn't win everywhere, though: GTZAN beat tracking and fine-grained structure boundaries still favor prior specialists.
Alex: But F1 on notes doesn't tell me the score is coherent. Did they measure coherence directly?
Sam: They did, and they flag that standard metrics miss it. Melody F1 here ignores octaves, and beat F1 doesn't capture tempo flips. So they count recordings that mix tempo levels. On the full-length osu2017 recordings, a strong recent beat tracker, Beat This!, mixes tempo levels in about 21 percent of songs. The distilled model does so in about one and a half percent. Applying a conventional beat decoder to the same Prober outputs also did worse than the anchor approach.
Alex: And octaves?
Sam: On 333 passages of RWC full melody, SheetSage1 mixes octaves in about 35 percent. The full decoder gets that to about 9 percent. Removing the interval dependence raises it to about 13. So the interval scoring helps, but most of the gain over SheetSage1 shows up either way. The distilled model sits in between, at about 12.
Alex: Does synthetic audio really transfer? A software-synthesized MIDI file sounds nothing like a real band.
Sam: Mostly, by their ablation. Adding synthetic data improves the beat, downbeat, chord and structure averages. Melody stays flat, and key slightly drops. The striking part: a model trained only on synthetic audio reached about 58 percent vocal melody F1, though its training audio contained no singing. They read that as suggesting simplified audio can transfer, not as proof.
Alex: What should make me cautious?
Sam: A few things. Some tempo and octave tests use Hooktheory data that overlaps training, which they acknowledge. Baselines run under differing setups. But the biggest limit is built into the design. Notes sit on a sixteenth-note grid, so triplets are approximated. One tempo anchor per song rules out big tempo changes. And the chord vocabulary stops at sevenths, so ninths can't be written. Training and evaluation are dominated by Western pop with a steady beat. They say plainly it isn't designed for classical or free-rhythm music.
Alex: So who should sit down with the full report, and where should they start?
Sam: Anyone in music information retrieval, or building symbolic data for music generation. Start with the structured decoding sections on rhythm and melody. Then the synthetic-data ablation, and the appendix on consistency diagnostics, which shows what the headline scores leave out.
Alex: And for everyone else, the line to carry around?
Sam: Right notes don't make a score. Consistency across the whole song does, and here it's taught through rules, then handed to one model.
Alex: A useful reminder that what you measure shapes what you get.