Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AI Music Production Copilot

Describe a vibe in plain English, get back a finished instrumental — MIDI, a mixed WAV, and separate stems. No DAW required.

ai-music-vibe "deep bass and intense hero music" --out-wav output/hero.wav
Generated: output/hero.mid
Mixed WAV: output/hero.wav
Stems:     output/stems/
Session:   output/hero.session.json
LLM patterns: yes

Composition params:
  BPM:     128
  Key:     A minor
  Rhythm:  film_pop
  Raga:    kalyani
  Style:   anthem
  Genre:   anthem

It specializes in South Indian film-pop, kuthu, gaana, and cinematic score — the theory layer knows Carnatic ragas and kuthu rhythm templates, not just Western pop chords.

How it works

An LLM writes the patterns; deterministic Python writes the music. The model never emits audio — it emits JSON note data, which is then validated against music theory, arranged into sections, and rendered by a real synth/sampler chain.

vibe (+ optional reference track)
  → LLM proposes params + patterns (drums, bass, chords, hook)
  → theory validator          # snap to raga/scale, clamp ranges
  → arrangement assembler     # intro / build / drop / outro per genre
  → hybrid render             # sample drums + 808 synth + SoundFont
  → auto-mix                  # per-genre gain presets, FX, stems
  → optional analyze → refine loop

Two properties fall out of this design:

  • The LLM is optional. No API key, and it silently falls back to the rule-based composer. ai-music-render never calls an LLM at all.
  • It's reproducible. The seed is derived from a hash of your vibe string, so the same words give the same track. Pass --seed to pin it explicitly.

Install

python3 -m venv venv
source venv/bin/activate
pip install -e ".[llm]"          # drop [llm] for the offline rule-based composer

cp .env.example .env             # add OPENAI_API_KEY or Azure credentials

Requires Python 3.9+. Rendering uses FluidSynth via a SoundFont — point at one with --soundfont path/to.sf2 if the bundled default isn't found.

The three commands

ai-music-vibe — generate from a description

# Straight generation
ai-music-vibe "kuthu banger" --out-midi output/track.mid --out-wav output/track.wav

# Match an existing track's tempo and energy
ai-music-vibe "dark trailer score" --reference path/to/song.mp3 --out-wav output/track.wav

# Generate → analyze the audio → let the LLM fix its own mix → repeat
ai-music-vibe "emotional background score" --iterations 3 --out-wav output/track.wav

--iterations is the interesting one. After each pass it renders the WAV, runs a spectral/tempo analysis on it, and feeds the critique back to the model ("sound is very bright — may need warmth in mix"), which rewrites the patterns.

ai-music-refine — nudge a finished track

Every generation drops a .session.json. Point at it and ask for a change in English:

ai-music-refine "make the drop harder and the hook brighter" \
  --session output/track.session.json

ai-music-render — your own patterns, zero AI

If you already have the notes, skip the model entirely and just use this as a renderer and mixer:

ai-music-render examples/tracks.phonk.json --out-wav output/phonk.wav

Writing a composition by hand

The track format has no fixed instrument slots — add as many as you want.

{
  "bpm": 112,
  "genre": "phonk",
  "tracks": [
    { "name": "kick",    "engine": "drums",         "hits": [{ "instrument": "kick", "step": 0, "velocity": 127 }] },
    { "name": "808",     "engine": "synth_808",     "layers": ["sub", "body", "grit"], "notes": [] },
    { "name": "cowbell", "engine": "synth_cowbell", "notes": [], "pan": 0.25 },
    { "name": "pad",     "engine": "soundfont",     "program": 89, "notes": [] },
    { "name": "vinyl",   "engine": "sample",        "sample": "loops/vinyl.wav", "loop": true, "gain": 0.15 }
  ]
}

Engines: drums, synth_808, synth_cowbell, synth, soundfont, sample

  • Percussion tracks use hits — { instrument, step, velocity } on a 16-step grid (set steps to change resolution). Any instrument name works; it resolves to a sample file if one exists, otherwise a built-in synth voice. kick, snare, clap, hihat_closed, hihat_open, tom_low, tom_mid, crash, ride, cowbell, perc, and shaker are built in.
  • Pitched tracks use notes — { midi, start_beat, duration_beats, velocity }. Patterns shorter than a section loop automatically.
  • Per-track pan, gain_db, and effects (reverb, distortion, compressor, lowpass, bright_shelf, transient) feed the mixer.

See examples/tracks.phonk.json for a complete file, and examples/composition.example.json for the older 4-pattern format (drums/bass/chords/hook), which is still accepted and auto-converted into tracks.

What the vibe parser understands

Without an API key, your words are matched against keyword profiles. These are worth knowing because they're what the fallback keys off:

Say something like You get BPM
cinematic, emotional, trailer, background film score, Kalyani, major 92
romantic, love, soft, ballad ballad, Bilahari, major 80
dark, moody, suspense dark score, Shanmukhapriya, minor 110
anthem, heroic, epic, victory anthem, Kalyani, major 128
gaana, street, chennai gaana, Mohanam, minor 130
dappankuthu, mass dappankuthu, Hamsadhwani 160
kuthu, banger, party, club kuthu, Hamsadhwani, minor 150
trap, hip hop, 808 trap-kuthu, Revati, minor 140

Modifiers stack on top: an explicit "128 bpm", a raga by name, "in F#", and slow/fast/bright/dark all override the profile. With an API key, the LLM picks these parameters instead and the keyword table is just the safety net.

Output

File What it is
track.mid MIDI, one instrument per channel
track.wav the auto-mixed stereo master
stems/ each track rendered on its own, for a real DAW
track.session.json params + patterns, the input to ai-music-refine

Development

pip install -e ".[dev]"
pytest

License

MIT

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages