2026 - Important Thesis - Spatial Music Generation
Continuous Tokenization from Transformer, Jyrki, Industry, AI Center
References 1
- 📍 2025 - Audio Compression, J. Alakuijala
- 2020 - DDSP: Differentiable Digital Signal Processing, ICLR
- 2026 - GDM, Learning the integral of a diffusion model, Sander Dieleman
- 2015 - BiternionNets: continuous head orientation from discrete labels
- 2025 - Who Invented Transformer Neural Networks?
- 1960 - A new approach to linear filtering and prediction problems, Kalman, R E
- 📍 2021 - A Mathematical Framework for Transformer Circuits, Anthropic
- 2026 - Do Value Vectors in Deep Layers Need Context from the Residual Stream, AI/ML
- Classic Prediction Models
- 2018 - Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset, Google Brain, Deepmind
- András Schiff
- 1979 - Gödel, Escher, Bach: an Eternal Golden Braid
Tokenization - from Audio to LLM
In multimodal generation, tokenization extends beyond text: images, speech, and music can be represented
either as discrete codebook indices, as in VQ-VAE, SoundStream, and EnCodec, or as continuous latent tokens,
as in patch-based Transformers, latent diffusion, and flow-matching models. Our work follows the latter
direction by treating music generation as continuous perceptual tokenization, where pitch, rhythm, timbre,
and phrase-level structure are preserved as continuous trajectories rather than being forced into a fixed
discrete vocabulary.
Generation
- prompting:
write in the most simple and dense way what is the substantial difference between flow matching and other diffusion approaches
Evaluation
| Context | Correct Term |
|---|---|
| Image generation | FID |
| Audio / music generation | FAD |
| General mathematical distance | Fréchet Distance / FD |
| Metric | Generated object | Ground-truth / reference source | What is compared | Loss / distance / score | Interpretation |
|---|---|---|---|---|---|
| FAD / FD-audio ↓ | Rendered WAV/audio generated by our Flow-Matching DiT and by external baselines | Held-out real Mozart recordings, or held-out Mozart MIDI rendered with the same soundfont / piano renderer | Distribution of generated audio embeddings vs. distribution of real Mozart audio embeddings | Fréchet Audio Distance / Fréchet Distance in audio embedding space | Measures whether the generated audio distribution is close to real Mozart audio at the corpus level. Lower means better audio realism and closer global acoustic distribution. |
| MERT embedding FD ↓ | Generated WAV/audio converted into MERT music embeddings | Held-out real Mozart audio converted into MERT embeddings | Distribution of generated music-semantic embeddings vs. distribution of Mozart reference embeddings | Fréchet Distance in MERT embedding space | Measures whether the generated samples match Mozart-like musical semantics, including melody, harmony, rhythm, timbre, and style-level structure. Lower means stronger Mozart-style alignment. |
| CLAP text-audio score ↑ | Generated WAV/audio from each text prompt | The input text prompt used to condition generation, e.g., “Mozart-style solo piano sonata in allegro” | Text embedding of the prompt vs. audio embedding of the generated sample | CLAP cosine similarity / text-audio alignment score | Measures whether the generated music follows the conditioning prompt. Higher means better prompt controllability and text-to-music alignment. |
| Spectral centroid / bandwidth / rolloff KL ↓ | Spectral statistics extracted from generated WAV/audio | Spectral statistics extracted from held-out real Mozart recordings or consistently rendered Mozart MIDI | Distribution of generated spectral features vs. distribution of real Mozart spectral features | KL divergence or distribution distance over spectral centroid, bandwidth, and rolloff | Measures whether the generated audio has Mozart-like timbre, brightness, frequency balance, and instrumental texture. Lower means closer spectral behavior. |
| Onset density distance ↓ | Onset curve extracted from generated MIDI/note trajectories, or from generated WAV using an onset detector | Onset curve extracted from held-out Mozart MIDI/note tables, or from real Mozart audio using the same detector | Generated onset activity over time vs. Mozart reference onset activity over time | L1/L2 distance, KL divergence, or Earth Mover’s Distance over onset-density distributions | Measures whether the generated music has Mozart-like rhythmic activity, note attack density, and phrase- |
Audio (Symphonic Music Generation)
Others
- Keith Jarrett - Over the Rainbow (Tokyo 1984) [Restored]
- 1851 - Franz Liszt - Campanella
- [2019 - Lang Lang – Bach: The Well-Tempered Clavier: Book 1, 1.Prelude C Major, BWV 846]
GPT’s underlying capabilities come from
| Learned Ability | How It Emerges |
|---|---|
| Grammar | Predicting the next token requires the model to learn syntactic and grammatical patterns. |
| Facts | Predicting text accurately requires the model to absorb factual and world knowledge from large-scale corpora. |
| Reasoning Patterns | Explanations, proofs, code, mathematical derivations, and problem-solving texts contain reusable reasoning structures. |
| Style Imitation | The training data contains many writing styles, such as papers, emails, code, dialogues, poems, and documentation. |
| Translation | Multilingual corpora contain cross-lingual correspondences between words, phrases, and meanings. |
| Coding | Code itself is a token sequence with syntax, semantics, libraries, and execution-like patterns. |
| Few-shot Learning | When the model sees examples in the context, it learns to imitate the task format and continue the pattern. |
RLHF
model generates answer
↓
reward model scores answer
↓
policy optimization updates model
↓
model becomes more aligned with human preferences
| Stage | Training Signal | What the Model Learns | Result |
|---|---|---|---|
| Pre-training | Next-token prediction on massive text | Language, facts, code, reasoning patterns, world knowledge | Base GPT |
| Supervised Fine-Tuning | Human-written instruction-answer pairs | How to answer instructions | Instruct model |
| Reward Modeling | Human preference comparisons | Which answers humans prefer | Preference scorer |
| RLHF / Preference Optimization | Reward model feedback | More helpful, harmless, instruction-following behavior | ChatGPT-style assistant |
| Safety / policy tuning | Refusal and safety examples | When to refuse, how to be careful | Safer deployed model |
| Tool / multimodal training, if applicable | Tool-use traces, image/audio/text data | Search, coding, image/audio understanding, tool calling | Modern assistant system |
Tokenization
| Stage | Earliest / canonical paper or model in modern AI/ML | Year | Authors / Organization | Token Type | What was introduced | Representative Methods |
|---|---|---|---|---|---|---|
| Rule-based tokenization | The Penn Treebank / rule-based corpus tokenization conventions | 1993 | Mitchell P. Marcus, Beatrice Santorini, Mary Ann Marcinkiewicz | Discrete hand-defined tokens | A widely adopted early NLP convention for splitting and normalizing text into linguistically interpretable word and punctuation tokens for corpus annotation. This is a canonical reference point for rule-based tokenization in modern NLP, although rule-based tokenization itself existed earlier. | Whitespace splitting, punctuation splitting, Penn Treebank tokenizer, WordPunct-style tokenizers |
| Dictionary-based segmentation | Japanese and Korean Voice Search / practical lexicon-based segmentation for non-whitespace languages | 2012 | Mike Schuster, Kaisuke Nakajima, Google | Discrete dictionary/statistical word tokens | A practical large-scale system-level reference for handling languages where token boundaries are not directly marked by whitespace. It reflects the move from universal whitespace tokenization to language-specific segmentation pipelines. | Jieba-style Chinese segmentation, MeCab-style Japanese segmentation, dictionary + HMM / statistical segmentation |
| Original Byte Pair Encoding | A New Algorithm for Data Compression | 1994 | Philip Gage | Discrete compression symbols | Introduced Byte Pair Encoding as a data-compression algorithm that repeatedly replaces the most frequent adjacent byte/symbol pair with a new symbol. It was not originally proposed for neural NLP, but later became the foundation of subword tokenization. | Original BPE compression |
| Subword tokenization | Neural Machine Translation of Rare Words with Subword Units | 2015 / 2016 | Rico Sennrich, Barry Haddow, Alexandra Birch | Discrete learned subword tokens | Adapted BPE to neural machine translation by representing rare and unknown words as sequences of subword units, reducing the out-of-vocabulary problem. | Subword BPE, BPE for NMT, later Transformer NMT tokenizers |
| WordPiece tokenization | Japanese and Korean Voice Search | 2012 | Mike Schuster, Kaisuke Nakajima, Google | Discrete learned subword tokens | Introduced WordPiece-style subword modeling in a Google speech/search system, later becoming the tokenizer family used in BERT-style models. | WordPiece, BERT tokenizer |
| Unigram Language Model tokenization | Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates | 2018 | Taku Kudo | Discrete probabilistic subword tokens | Proposed a unigram language-model-based subword segmentation algorithm and subword regularization, allowing multiple possible tokenizations to be sampled during training. | SentencePiece Unigram, subword regularization |
| Language-agnostic raw-text tokenization | SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing | 2018 | Taku Kudo, John Richardson | Discrete raw-text subword tokens | Removed the need for external pre-tokenization by training directly from raw text and treating whitespace as a normal symbol. | SentencePiece BPE, SentencePiece Unigram |
| Byte-level BPE | Language Models are Unsupervised Multitask Learners / GPT-2 tokenizer | 2019 | Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, OpenAI | Discrete byte-level subword tokens | Popularized byte-level BPE for large language models, converting text to UTF-8 byte sequences and applying BPE to avoid unknown tokens across arbitrary text. | GPT-2 tokenizer, RoBERTa byte-level BPE, GPT-style byte-level BPE |
| Systems-optimized tokenization | tiktoken | 2022 | OpenAI | Discrete production token IDs | Reframed tokenization as a high-throughput systems component for large-scale LLM training, inference, and token counting, using a fast BPE implementation. | tiktoken-style optimized BPE tokenizers, Hugging Face fast tokenizers, production LLM tokenizers |
| Learned discrete neural tokenization | Neural Discrete Representation Learning / VQ-VAE | 2017 | Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu | Discrete learned latent tokens | Introduced VQ-VAE, where a neural encoder maps continuous inputs to discrete codebook entries. This extended tokenization from text into learned discrete representations for images, audio, video, and speech. | VQ-VAE, VQGAN, DALL-E-style image tokenizers |
| High-resolution image codebook tokenization | Taming Transformers for High-Resolution Image Synthesis / VQGAN | 2020 / 2021 | Patrick Esser, Robin Rombach, Björn Ommer | Discrete visual codebook tokens | Combined a convolutional VQGAN codebook with an autoregressive Transformer, making high-resolution image generation possible through learned discrete visual tokens. | VQGAN, visual codebook tokenizers, autoregressive image token models |
| Neural audio codec tokenization | SoundStream: An End-to-End Neural Audio Codec | 2021 | Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi, Google | Discrete neural audio codec tokens | Introduced an end-to-end neural audio codec using residual vector quantization to compress speech, music, and general audio into quantized embeddings. | SoundStream, residual vector quantization audio codecs |
| High-fidelity audio codec tokenization | High Fidelity Neural Audio Compression / EnCodec | 2022 | Alexandre Défossez, Jade Copet, Gabriel Synnaeve, Yossi Adi, Meta AI | Discrete neural audio codec tokens | Introduced a high-fidelity neural audio codec with quantized latent space, later used as the tokenization backend for many text-to-audio and text-to-music systems. | EnCodec, AudioCraft / MusicGen audio tokens |
| Hierarchical neural audio tokenization | AudioLM: A Language Modeling Approach to Audio Generation | 2022 | Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, Neil Zeghidour, Google | Mostly discrete multi-level audio tokens | Combined semantic tokens for long-range structure with acoustic codec tokens for high-fidelity synthesis, establishing a hierarchical token view for audio generation. | AudioLM, MusicLM-style pipelines, semantic tokens + acoustic tokens |
| Patch tokenization | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale / Vision Transformer | 2020 / 2021 | Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby, Google | Continuous vector tokens | Represented an image as a sequence of fixed-size patch embeddings, turning visual input into Transformer-compatible continuous tokens. | ViT, patch embeddings, image patch tokens |
| Masked patch / latent tokenization | Masked Autoencoders Are Scalable Vision Learners | 2021 / 2022 | Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, Meta AI | Continuous vector tokens | Showed that masked image patches can be reconstructed from latent representations, strengthening the view of patches as scalable continuous tokens for vision. | MAE, masked patch modeling |
| Latent diffusion tokenization | High-Resolution Image Synthesis with Latent Diffusion Models | 2021 / 2022 | Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer | Continuous latent tokens | Moved diffusion from pixel space into a learned continuous latent space, making high-resolution generation more efficient. | Latent Diffusion Models, Stable Diffusion-style latent tokenization |
| Diffusion Transformer latent tokens | Scalable Diffusion Models with Transformers / DiT | 2022 / 2023 | William Peebles, Saining Xie | Continuous latent patch tokens | Replaced the U-Net backbone in latent diffusion with a Transformer operating on latent patches, connecting patch tokenization with generative diffusion. | DiT, latent patch tokens, Transformer diffusion |
| Tokenizer-free modeling | CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation | 2021 / 2022 | Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, Google | Low-level character tokens without explicit tokenizer vocabulary | Proposed a neural encoder that operates directly on character sequences without a fixed tokenizer vocabulary, reducing dependence on external tokenization. | CANINE, character-level tokenization-free encoders |
| Tokenizer-free byte modeling | ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models | 2021 / 2022 | Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, Google | Low-level byte tokens | Showed that a standard Transformer can operate directly on UTF-8 bytes, removing subword tokenization while remaining competitive with token-based models. | ByT5, byte-level Transformers |
| End-to-end learned tokenization | Charformer: Fast Character Transformers via Gradient-based Subword Tokenization | 2021 | Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, Donald Metzler, Google | Learned latent subword representations from characters | Learned subword-like representations end-to-end inside the model through gradient-based subword tokenization, instead of using a fixed external tokenizer. | Charformer, GBST |
| Long-context byte / tokenizer-free modeling | MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers | 2023 | Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis, Meta AI | Byte-level tokens with multi-scale patches | Made byte-level modeling more scalable by using a multi-scale Transformer over byte patches, supporting long sequences such as books, audio files, and images. | MEGABYTE, byte patches, multiscale Transformers |
Reading Mel-Spectrogram
| Spectrogram Region | Likely Musical Meaning |
|---|---|
| 20–100 Hz | Sub-bass, bass drum, organ, double bass lowest notes |
| 100–300 Hz | Bass, cello, low piano, male voice body |
| 300–800 Hz | Warmth, lower harmonics, viola/cello/woodwinds |
| 800 Hz–2 kHz | Main melodic body, speech/song clarity, many orchestral harmonics |
| 2–5 kHz | Presence, brightness, attack, vocal intelligibility |
| 5–10 kHz | Air, bow noise, breath, cymbals, piano attack, brilliance |
| 10–16 kHz | Very high air/noise/detail, often weak in classical recordings |
Approximate Frequency Ranges of Classical Instruments
| Instrument Class | Instrument | Approximate Fundamental Range | Important Harmonic / Brightness Range | Typical Spectrogram Location |
|---|---|---|---|---|
| Strings | Double Bass | 41–300 Hz | 300 Hz–3 kHz | Low frequency body with upper harmonic lines |
| Strings | Cello | 65–1,000 Hz | 200 Hz–5 kHz | Low-to-mid frequency melodic body |
| Strings | Viola | 130–1,300 Hz | 300 Hz–6 kHz | Mid-frequency warm harmonic texture |
| Strings | Violin | 196–3,500 Hz | 1–10 kHz | Mid-to-high melodic and bright harmonic lines |
| Strings | Harp | 30–3,000 Hz | 500 Hz–8 kHz | Wide range, plucked vertical transients |
| Woodwinds | Contrabassoon | 29–300 Hz | 200 Hz–3 kHz | Very low dark body |
| Woodwinds | Bassoon | 58–700 Hz | 300 Hz–5 kHz | Low-mid nasal harmonic structure |
| Woodwinds | Clarinet | 147–1,800 Hz | 500 Hz–8 kHz | Mid-frequency smooth tone, strong odd harmonics |
| Woodwinds | Oboe | 247–1,500 Hz | 1–8 kHz | Bright mid-high penetrating tone |
| Woodwinds | Flute | 262–2,100 Hz | 1–10 kHz | High, airy, less dense harmonic body |
| Woodwinds | Piccolo | 587–4,000 Hz | 2–12 kHz | Very high bright energy |
| Brass | Tuba | 40–400 Hz | 200 Hz–4 kHz | Low-frequency power and brass harmonics |
| Brass | Trombone | 80–600 Hz | 300 Hz–6 kHz | Low-mid strong brass energy |
| Brass | French Horn | 65–1,000 Hz | 300 Hz–6 kHz | Warm mid-frequency harmonic mass |
| Brass | Trumpet | 165–1,000 Hz | 1–10 kHz | Bright, strong upper harmonics |
| Percussion | Timpani | 60–300 Hz | 200 Hz–3 kHz | Low pitched drum resonance |
| Percussion | Bass Drum | 30–150 Hz | 100 Hz–5 kHz | Low boom plus broadband attack |
| Percussion | Snare Drum | 150–500 Hz | 1–10 kHz | Broadband vertical transient |
| Percussion | Cymbals | No stable pitch | 2–16 kHz | High-frequency noisy shimmer |
| Percussion | Triangle | 2–10 kHz | 5–18 kHz | Very high bright narrow energy |
| Keyboard | Piano | 27.5–4,186 Hz | 100 Hz–10 kHz | Full-range vertical attacks plus harmonic decay |
| Keyboard | Organ | 16–8,000 Hz | 100 Hz–12 kHz | Sustained wide-band harmonic layers |
| Voice-like Classical Timbre | Choir / Orchestra Blend | 80–4,000 Hz | 500 Hz–10 kHz | Dense mid-frequency harmonic texture |
Approximate Frequency Ranges of Human Singing Voices
| Voice Type | Approximate Fundamental Range | Important Harmonic / Formant Range | Spectrogram Interpretation |
|---|---|---|---|
| Bass | 80–330 Hz | 300 Hz–4 kHz | Low male voice body with harmonic lines |
| Baritone | 100–400 Hz | 300 Hz–5 kHz | Low-mid male voice, strong body |
| Tenor | 130–520 Hz | 500 Hz–6 kHz | Higher male voice, clearer upper harmonics |
| Alto / Contralto | 165–700 Hz | 500 Hz–6 kHz | Low female voice, warm mid-frequency body |
| Mezzo-soprano | 220–900 Hz | 700 Hz–8 kHz | Mid-high female voice, strong presence |
| Soprano | 260–1,200 Hz | 1–10 kHz | High female voice, bright upper harmonic energy |
| Children’s Voice | 250–1,000 Hz | 1–8 kHz | High fundamental, light upper harmonics |
| Spoken Male Voice | 85–180 Hz | 300 Hz–4 kHz | Lower speech pitch, intelligibility around 1–4 kHz |
| Spoken Female Voice | 165–255 Hz | 500 Hz–5 kHz | Higher speech pitch, intelligibility around 1–5 kHz |
Tools
| Library / Repository | Author(s) + Proposed Year |
|---|---|
librosa | McFee et al., 2015 |
torchaudio | PyTorch Audio team / Hwang et al., 2023 |
descript-audio-codec / DAC | Kumar, Seetharaman, Luebs, I. Kumar, K. Kumar, 2023 |
DAC-JAX | David Braun, 2024 |
encodec | Défossez, Copet, Synnaeve, Adi, 2022 |
audiocraft / MusicGen | Copet et al., 2023 |
stable-audio-tools | Stability AI, 2023–present |
ddsp | Engel et al., 2020 |
Gemma 2026, Best practical choice
| Rank | Model Link / ID | Use in Your Project | Why |
|---|---|---|---|
| 1 | google/gemma-4-E2B-it | Default Music LLM prompt parser | Best balance. Small enough for A100 experiments, supports structured prompting, native system prompt, Apache 2.0, and does not steal too much memory from AiT. |
| 2 | google/gemma-4-E4B-it | Stronger prompt parser / optional ablation | Better reasoning than E2B, still lightweight, but larger. Use after the E2B pipeline works. |
| 3 | google/gemma-4-12B-it | Offline high-quality parser only | Stronger, but unnecessary for fixed-schema JSON parsing. Good for offline auto-labeling / prompt expansion, not ideal as default during audio training. |
| 4 | Qwen/Qwen2.5-3B-Instruct | Non-Gemma baseline | Strong for structured outputs and JSON. Qwen’s model card explicitly highlights improvements in structured data and JSON generation. (Hugging Face) |
| 5 | microsoft/Phi-3.5-mini-instruct | Small LLM baseline | MIT license and 3.8B parameters; tested on A100 according to the model card. Good baseline, but less directly aligned with your Gemma plan. (Hugging Face) |