Describe a vibe in plain English, get back a finished instrumental — MIDI, a mixed WAV, and separate stems. No DAW required.
ai-music-vibe "deep bass and intense hero music" --out-wav output/hero.wavGenerated: output/hero.mid
Mixed WAV: output/hero.wav
Stems: output/stems/
Session: output/hero.session.json
LLM patterns: yes
Composition params:
BPM: 128
Key: A minor
Rhythm: film_pop
Raga: kalyani
Style: anthem
Genre: anthem
It specializes in South Indian film-pop, kuthu, gaana, and cinematic score — the theory layer knows Carnatic ragas and kuthu rhythm templates, not just Western pop chords.
An LLM writes the patterns; deterministic Python writes the music. The model never emits audio — it emits JSON note data, which is then validated against music theory, arranged into sections, and rendered by a real synth/sampler chain.
vibe (+ optional reference track)
→ LLM proposes params + patterns (drums, bass, chords, hook)
→ theory validator # snap to raga/scale, clamp ranges
→ arrangement assembler # intro / build / drop / outro per genre
→ hybrid render # sample drums + 808 synth + SoundFont
→ auto-mix # per-genre gain presets, FX, stems
→ optional analyze → refine loop
Two properties fall out of this design:
- The LLM is optional. No API key, and it silently falls back to the rule-based composer.
ai-music-rendernever calls an LLM at all. - It's reproducible. The seed is derived from a hash of your vibe string, so the same words give the same track. Pass
--seedto pin it explicitly.
python3 -m venv venv
source venv/bin/activate
pip install -e ".[llm]" # drop [llm] for the offline rule-based composer
cp .env.example .env # add OPENAI_API_KEY or Azure credentialsRequires Python 3.9+. Rendering uses FluidSynth via a SoundFont — point at one with --soundfont path/to.sf2 if the bundled default isn't found.
# Straight generation
ai-music-vibe "kuthu banger" --out-midi output/track.mid --out-wav output/track.wav
# Match an existing track's tempo and energy
ai-music-vibe "dark trailer score" --reference path/to/song.mp3 --out-wav output/track.wav
# Generate → analyze the audio → let the LLM fix its own mix → repeat
ai-music-vibe "emotional background score" --iterations 3 --out-wav output/track.wav--iterations is the interesting one. After each pass it renders the WAV, runs a spectral/tempo analysis on it, and feeds the critique back to the model ("sound is very bright — may need warmth in mix"), which rewrites the patterns.
Every generation drops a .session.json. Point at it and ask for a change in English:
ai-music-refine "make the drop harder and the hook brighter" \
--session output/track.session.jsonIf you already have the notes, skip the model entirely and just use this as a renderer and mixer:
ai-music-render examples/tracks.phonk.json --out-wav output/phonk.wavThe track format has no fixed instrument slots — add as many as you want.
{
"bpm": 112,
"genre": "phonk",
"tracks": [
{ "name": "kick", "engine": "drums", "hits": [{ "instrument": "kick", "step": 0, "velocity": 127 }] },
{ "name": "808", "engine": "synth_808", "layers": ["sub", "body", "grit"], "notes": [] },
{ "name": "cowbell", "engine": "synth_cowbell", "notes": [], "pan": 0.25 },
{ "name": "pad", "engine": "soundfont", "program": 89, "notes": [] },
{ "name": "vinyl", "engine": "sample", "sample": "loops/vinyl.wav", "loop": true, "gain": 0.15 }
]
}Engines: drums, synth_808, synth_cowbell, synth, soundfont, sample
- Percussion tracks use
hits—{ instrument, step, velocity }on a 16-step grid (setstepsto change resolution). Anyinstrumentname works; it resolves to a sample file if one exists, otherwise a built-in synth voice.kick,snare,clap,hihat_closed,hihat_open,tom_low,tom_mid,crash,ride,cowbell,perc, andshakerare built in. - Pitched tracks use
notes—{ midi, start_beat, duration_beats, velocity }. Patterns shorter than a section loop automatically. - Per-track
pan,gain_db, andeffects(reverb,distortion,compressor,lowpass,bright_shelf,transient) feed the mixer.
See examples/tracks.phonk.json for a complete file, and examples/composition.example.json for the older 4-pattern format (drums/bass/chords/hook), which is still accepted and auto-converted into tracks.
Without an API key, your words are matched against keyword profiles. These are worth knowing because they're what the fallback keys off:
| Say something like | You get | BPM |
|---|---|---|
cinematic, emotional, trailer, background |
film score, Kalyani, major | 92 |
romantic, love, soft, ballad |
ballad, Bilahari, major | 80 |
dark, moody, suspense |
dark score, Shanmukhapriya, minor | 110 |
anthem, heroic, epic, victory |
anthem, Kalyani, major | 128 |
gaana, street, chennai |
gaana, Mohanam, minor | 130 |
dappankuthu, mass |
dappankuthu, Hamsadhwani | 160 |
kuthu, banger, party, club |
kuthu, Hamsadhwani, minor | 150 |
trap, hip hop, 808 |
trap-kuthu, Revati, minor | 140 |
Modifiers stack on top: an explicit "128 bpm", a raga by name, "in F#", and slow/fast/bright/dark all override the profile. With an API key, the LLM picks these parameters instead and the keyword table is just the safety net.
| File | What it is |
|---|---|
track.mid |
MIDI, one instrument per channel |
track.wav |
the auto-mixed stereo master |
stems/ |
each track rendered on its own, for a real DAW |
track.session.json |
params + patterns, the input to ai-music-refine |
pip install -e ".[dev]"
pytestMIT