2026 - Important Thesis - Spatial Music Generation

Continuous Tokenization from Transformer, Jyrki, Industry, AI Center


References 1









Tokenization - from Audio to LLM

In multimodal generation, tokenization extends beyond text: images, speech, and music can be represented
either as discrete codebook indices, as in VQ-VAE, SoundStream, and EnCodec, or as continuous latent tokens,
as in patch-based Transformers, latent diffusion, and flow-matching models. Our work follows the latter
direction by treating music generation as continuous perceptual tokenization, where pitch, rhythm, timbre,
and phrase-level structure are preserved as continuous trajectories rather than being forced into a fixed
discrete vocabulary.






Generation

  • prompting:
write in the most simple and dense way what is the substantial difference between flow matching and other diffusion approaches













Evaluation

Context Correct Term
Image generation FID
Audio / music generation FAD
General mathematical distance Fréchet Distance / FD


Metric Generated object Ground-truth / reference source What is compared Loss / distance / score Interpretation
FAD / FD-audio ↓ Rendered WAV/audio generated by our Flow-Matching DiT and by external baselines Held-out real Mozart recordings, or held-out Mozart MIDI rendered with the same soundfont / piano renderer Distribution of generated audio embeddings vs. distribution of real Mozart audio embeddings Fréchet Audio Distance / Fréchet Distance in audio embedding space Measures whether the generated audio distribution is close to real Mozart audio at the corpus level. Lower means better audio realism and closer global acoustic distribution.
MERT embedding FD ↓ Generated WAV/audio converted into MERT music embeddings Held-out real Mozart audio converted into MERT embeddings Distribution of generated music-semantic embeddings vs. distribution of Mozart reference embeddings Fréchet Distance in MERT embedding space Measures whether the generated samples match Mozart-like musical semantics, including melody, harmony, rhythm, timbre, and style-level structure. Lower means stronger Mozart-style alignment.
CLAP text-audio score ↑ Generated WAV/audio from each text prompt The input text prompt used to condition generation, e.g., “Mozart-style solo piano sonata in allegro” Text embedding of the prompt vs. audio embedding of the generated sample CLAP cosine similarity / text-audio alignment score Measures whether the generated music follows the conditioning prompt. Higher means better prompt controllability and text-to-music alignment.
Spectral centroid / bandwidth / rolloff KL ↓ Spectral statistics extracted from generated WAV/audio Spectral statistics extracted from held-out real Mozart recordings or consistently rendered Mozart MIDI Distribution of generated spectral features vs. distribution of real Mozart spectral features KL divergence or distribution distance over spectral centroid, bandwidth, and rolloff Measures whether the generated audio has Mozart-like timbre, brightness, frequency balance, and instrumental texture. Lower means closer spectral behavior.
Onset density distance ↓ Onset curve extracted from generated MIDI/note trajectories, or from generated WAV using an onset detector Onset curve extracted from held-out Mozart MIDI/note tables, or from real Mozart audio using the same detector Generated onset activity over time vs. Mozart reference onset activity over time L1/L2 distance, KL divergence, or Earth Mover’s Distance over onset-density distributions Measures whether the generated music has Mozart-like rhythmic activity, note attack density, and phrase-









Audio (Symphonic Music Generation)











Others











GPT’s underlying capabilities come from

Learned Ability How It Emerges
Grammar Predicting the next token requires the model to learn syntactic and grammatical patterns.
Facts Predicting text accurately requires the model to absorb factual and world knowledge from large-scale corpora.
Reasoning Patterns Explanations, proofs, code, mathematical derivations, and problem-solving texts contain reusable reasoning structures.
Style Imitation The training data contains many writing styles, such as papers, emails, code, dialogues, poems, and documentation.
Translation Multilingual corpora contain cross-lingual correspondences between words, phrases, and meanings.
Coding Code itself is a token sequence with syntax, semantics, libraries, and execution-like patterns.
Few-shot Learning When the model sees examples in the context, it learns to imitate the task format and continue the pattern.


RLHF

model generates answer
↓
reward model scores answer
↓
policy optimization updates model
↓
model becomes more aligned with human preferences


Stage Training Signal What the Model Learns Result
Pre-training Next-token prediction on massive text Language, facts, code, reasoning patterns, world knowledge Base GPT
Supervised Fine-Tuning Human-written instruction-answer pairs How to answer instructions Instruct model
Reward Modeling Human preference comparisons Which answers humans prefer Preference scorer
RLHF / Preference Optimization Reward model feedback More helpful, harmless, instruction-following behavior ChatGPT-style assistant
Safety / policy tuning Refusal and safety examples When to refuse, how to be careful Safer deployed model
Tool / multimodal training, if applicable Tool-use traces, image/audio/text data Search, coding, image/audio understanding, tool calling Modern assistant system





















Tokenization

Stage Earliest / canonical paper or model in modern AI/ML Year Authors / Organization Token Type What was introduced Representative Methods
Rule-based tokenization The Penn Treebank / rule-based corpus tokenization conventions 1993 Mitchell P. Marcus, Beatrice Santorini, Mary Ann Marcinkiewicz Discrete hand-defined tokens A widely adopted early NLP convention for splitting and normalizing text into linguistically interpretable word and punctuation tokens for corpus annotation. This is a canonical reference point for rule-based tokenization in modern NLP, although rule-based tokenization itself existed earlier. Whitespace splitting, punctuation splitting, Penn Treebank tokenizer, WordPunct-style tokenizers
Dictionary-based segmentation Japanese and Korean Voice Search / practical lexicon-based segmentation for non-whitespace languages 2012 Mike Schuster, Kaisuke Nakajima, Google Discrete dictionary/statistical word tokens A practical large-scale system-level reference for handling languages where token boundaries are not directly marked by whitespace. It reflects the move from universal whitespace tokenization to language-specific segmentation pipelines. Jieba-style Chinese segmentation, MeCab-style Japanese segmentation, dictionary + HMM / statistical segmentation
Original Byte Pair Encoding A New Algorithm for Data Compression 1994 Philip Gage Discrete compression symbols Introduced Byte Pair Encoding as a data-compression algorithm that repeatedly replaces the most frequent adjacent byte/symbol pair with a new symbol. It was not originally proposed for neural NLP, but later became the foundation of subword tokenization. Original BPE compression
Subword tokenization Neural Machine Translation of Rare Words with Subword Units 2015 / 2016 Rico Sennrich, Barry Haddow, Alexandra Birch Discrete learned subword tokens Adapted BPE to neural machine translation by representing rare and unknown words as sequences of subword units, reducing the out-of-vocabulary problem. Subword BPE, BPE for NMT, later Transformer NMT tokenizers
WordPiece tokenization Japanese and Korean Voice Search 2012 Mike Schuster, Kaisuke Nakajima, Google Discrete learned subword tokens Introduced WordPiece-style subword modeling in a Google speech/search system, later becoming the tokenizer family used in BERT-style models. WordPiece, BERT tokenizer
Unigram Language Model tokenization Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates 2018 Taku Kudo Discrete probabilistic subword tokens Proposed a unigram language-model-based subword segmentation algorithm and subword regularization, allowing multiple possible tokenizations to be sampled during training. SentencePiece Unigram, subword regularization
Language-agnostic raw-text tokenization SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing 2018 Taku Kudo, John Richardson Discrete raw-text subword tokens Removed the need for external pre-tokenization by training directly from raw text and treating whitespace as a normal symbol. SentencePiece BPE, SentencePiece Unigram
Byte-level BPE Language Models are Unsupervised Multitask Learners / GPT-2 tokenizer 2019 Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, OpenAI Discrete byte-level subword tokens Popularized byte-level BPE for large language models, converting text to UTF-8 byte sequences and applying BPE to avoid unknown tokens across arbitrary text. GPT-2 tokenizer, RoBERTa byte-level BPE, GPT-style byte-level BPE
Systems-optimized tokenization tiktoken 2022 OpenAI Discrete production token IDs Reframed tokenization as a high-throughput systems component for large-scale LLM training, inference, and token counting, using a fast BPE implementation. tiktoken-style optimized BPE tokenizers, Hugging Face fast tokenizers, production LLM tokenizers
Learned discrete neural tokenization Neural Discrete Representation Learning / VQ-VAE 2017 Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu Discrete learned latent tokens Introduced VQ-VAE, where a neural encoder maps continuous inputs to discrete codebook entries. This extended tokenization from text into learned discrete representations for images, audio, video, and speech. VQ-VAE, VQGAN, DALL-E-style image tokenizers
High-resolution image codebook tokenization Taming Transformers for High-Resolution Image Synthesis / VQGAN 2020 / 2021 Patrick Esser, Robin Rombach, Björn Ommer Discrete visual codebook tokens Combined a convolutional VQGAN codebook with an autoregressive Transformer, making high-resolution image generation possible through learned discrete visual tokens. VQGAN, visual codebook tokenizers, autoregressive image token models
Neural audio codec tokenization SoundStream: An End-to-End Neural Audio Codec 2021 Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi, Google Discrete neural audio codec tokens Introduced an end-to-end neural audio codec using residual vector quantization to compress speech, music, and general audio into quantized embeddings. SoundStream, residual vector quantization audio codecs
High-fidelity audio codec tokenization High Fidelity Neural Audio Compression / EnCodec 2022 Alexandre Défossez, Jade Copet, Gabriel Synnaeve, Yossi Adi, Meta AI Discrete neural audio codec tokens Introduced a high-fidelity neural audio codec with quantized latent space, later used as the tokenization backend for many text-to-audio and text-to-music systems. EnCodec, AudioCraft / MusicGen audio tokens
Hierarchical neural audio tokenization AudioLM: A Language Modeling Approach to Audio Generation 2022 Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, Neil Zeghidour, Google Mostly discrete multi-level audio tokens Combined semantic tokens for long-range structure with acoustic codec tokens for high-fidelity synthesis, establishing a hierarchical token view for audio generation. AudioLM, MusicLM-style pipelines, semantic tokens + acoustic tokens
Patch tokenization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale / Vision Transformer 2020 / 2021 Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby, Google Continuous vector tokens Represented an image as a sequence of fixed-size patch embeddings, turning visual input into Transformer-compatible continuous tokens. ViT, patch embeddings, image patch tokens
Masked patch / latent tokenization Masked Autoencoders Are Scalable Vision Learners 2021 / 2022 Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, Meta AI Continuous vector tokens Showed that masked image patches can be reconstructed from latent representations, strengthening the view of patches as scalable continuous tokens for vision. MAE, masked patch modeling
Latent diffusion tokenization High-Resolution Image Synthesis with Latent Diffusion Models 2021 / 2022 Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer Continuous latent tokens Moved diffusion from pixel space into a learned continuous latent space, making high-resolution generation more efficient. Latent Diffusion Models, Stable Diffusion-style latent tokenization
Diffusion Transformer latent tokens Scalable Diffusion Models with Transformers / DiT 2022 / 2023 William Peebles, Saining Xie Continuous latent patch tokens Replaced the U-Net backbone in latent diffusion with a Transformer operating on latent patches, connecting patch tokenization with generative diffusion. DiT, latent patch tokens, Transformer diffusion
Tokenizer-free modeling CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation 2021 / 2022 Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, Google Low-level character tokens without explicit tokenizer vocabulary Proposed a neural encoder that operates directly on character sequences without a fixed tokenizer vocabulary, reducing dependence on external tokenization. CANINE, character-level tokenization-free encoders
Tokenizer-free byte modeling ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models 2021 / 2022 Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, Google Low-level byte tokens Showed that a standard Transformer can operate directly on UTF-8 bytes, removing subword tokenization while remaining competitive with token-based models. ByT5, byte-level Transformers
End-to-end learned tokenization Charformer: Fast Character Transformers via Gradient-based Subword Tokenization 2021 Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, Donald Metzler, Google Learned latent subword representations from characters Learned subword-like representations end-to-end inside the model through gradient-based subword tokenization, instead of using a fixed external tokenizer. Charformer, GBST
Long-context byte / tokenizer-free modeling MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers 2023 Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis, Meta AI Byte-level tokens with multi-scale patches Made byte-level modeling more scalable by using a multi-scale Transformer over byte patches, supporting long sequences such as books, audio files, and images. MEGABYTE, byte patches, multiscale Transformers































Reading Mel-Spectrogram

Spectrogram Region Likely Musical Meaning
20–100 Hz Sub-bass, bass drum, organ, double bass lowest notes
100–300 Hz Bass, cello, low piano, male voice body
300–800 Hz Warmth, lower harmonics, viola/cello/woodwinds
800 Hz–2 kHz Main melodic body, speech/song clarity, many orchestral harmonics
2–5 kHz Presence, brightness, attack, vocal intelligibility
5–10 kHz Air, bow noise, breath, cymbals, piano attack, brilliance
10–16 kHz Very high air/noise/detail, often weak in classical recordings





Approximate Frequency Ranges of Classical Instruments

Instrument Class Instrument Approximate Fundamental Range Important Harmonic / Brightness Range Typical Spectrogram Location
Strings Double Bass 41–300 Hz 300 Hz–3 kHz Low frequency body with upper harmonic lines
Strings Cello 65–1,000 Hz 200 Hz–5 kHz Low-to-mid frequency melodic body
Strings Viola 130–1,300 Hz 300 Hz–6 kHz Mid-frequency warm harmonic texture
Strings Violin 196–3,500 Hz 1–10 kHz Mid-to-high melodic and bright harmonic lines
Strings Harp 30–3,000 Hz 500 Hz–8 kHz Wide range, plucked vertical transients
Woodwinds Contrabassoon 29–300 Hz 200 Hz–3 kHz Very low dark body
Woodwinds Bassoon 58–700 Hz 300 Hz–5 kHz Low-mid nasal harmonic structure
Woodwinds Clarinet 147–1,800 Hz 500 Hz–8 kHz Mid-frequency smooth tone, strong odd harmonics
Woodwinds Oboe 247–1,500 Hz 1–8 kHz Bright mid-high penetrating tone
Woodwinds Flute 262–2,100 Hz 1–10 kHz High, airy, less dense harmonic body
Woodwinds Piccolo 587–4,000 Hz 2–12 kHz Very high bright energy
Brass Tuba 40–400 Hz 200 Hz–4 kHz Low-frequency power and brass harmonics
Brass Trombone 80–600 Hz 300 Hz–6 kHz Low-mid strong brass energy
Brass French Horn 65–1,000 Hz 300 Hz–6 kHz Warm mid-frequency harmonic mass
Brass Trumpet 165–1,000 Hz 1–10 kHz Bright, strong upper harmonics
Percussion Timpani 60–300 Hz 200 Hz–3 kHz Low pitched drum resonance
Percussion Bass Drum 30–150 Hz 100 Hz–5 kHz Low boom plus broadband attack
Percussion Snare Drum 150–500 Hz 1–10 kHz Broadband vertical transient
Percussion Cymbals No stable pitch 2–16 kHz High-frequency noisy shimmer
Percussion Triangle 2–10 kHz 5–18 kHz Very high bright narrow energy
Keyboard Piano 27.5–4,186 Hz 100 Hz–10 kHz Full-range vertical attacks plus harmonic decay
Keyboard Organ 16–8,000 Hz 100 Hz–12 kHz Sustained wide-band harmonic layers
Voice-like Classical Timbre Choir / Orchestra Blend 80–4,000 Hz 500 Hz–10 kHz Dense mid-frequency harmonic texture


Approximate Frequency Ranges of Human Singing Voices

Voice Type Approximate Fundamental Range Important Harmonic / Formant Range Spectrogram Interpretation
Bass 80–330 Hz 300 Hz–4 kHz Low male voice body with harmonic lines
Baritone 100–400 Hz 300 Hz–5 kHz Low-mid male voice, strong body
Tenor 130–520 Hz 500 Hz–6 kHz Higher male voice, clearer upper harmonics
Alto / Contralto 165–700 Hz 500 Hz–6 kHz Low female voice, warm mid-frequency body
Mezzo-soprano 220–900 Hz 700 Hz–8 kHz Mid-high female voice, strong presence
Soprano 260–1,200 Hz 1–10 kHz High female voice, bright upper harmonic energy
Children’s Voice 250–1,000 Hz 1–8 kHz High fundamental, light upper harmonics
Spoken Male Voice 85–180 Hz 300 Hz–4 kHz Lower speech pitch, intelligibility around 1–4 kHz
Spoken Female Voice 165–255 Hz 500 Hz–5 kHz Higher speech pitch, intelligibility around 1–5 kHz
















Tools

Library / Repository Author(s) + Proposed Year
librosa McFee et al., 2015
torchaudio PyTorch Audio team / Hwang et al., 2023
descript-audio-codec / DAC Kumar, Seetharaman, Luebs, I. Kumar, K. Kumar, 2023
DAC-JAX David Braun, 2024
encodec Défossez, Copet, Synnaeve, Adi, 2022
audiocraft / MusicGen Copet et al., 2023
stable-audio-tools Stability AI, 2023–present
ddsp Engel et al., 2020








Gemma 2026, Best practical choice

Rank Model Link / ID Use in Your Project Why
1 google/gemma-4-E2B-it Default Music LLM prompt parser Best balance. Small enough for A100 experiments, supports structured prompting, native system prompt, Apache 2.0, and does not steal too much memory from AiT.
2 google/gemma-4-E4B-it Stronger prompt parser / optional ablation Better reasoning than E2B, still lightweight, but larger. Use after the E2B pipeline works.
3 google/gemma-4-12B-it Offline high-quality parser only Stronger, but unnecessary for fixed-schema JSON parsing. Good for offline auto-labeling / prompt expansion, not ideal as default during audio training.
4 Qwen/Qwen2.5-3B-Instruct Non-Gemma baseline Strong for structured outputs and JSON. Qwen’s model card explicitly highlights improvements in structured data and JSON generation. (Hugging Face)
5 microsoft/Phi-3.5-mini-instruct Small LLM baseline MIT license and 3.8B parameters; tested on A100 according to the model card. Good baseline, but less directly aligned with your Gemma plan. (Hugging Face)























References 2



References