Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Juno

Work In Progress

An offline, self-hosted AI music workstation with a Suno-style workflow, powered locally by ACE-Step 1.5 XL models. One Docker image runs the ACE-Step API server and the Juno web app; everything — model weights, generated audio, uploads, and your library — stays on your machine.

Sacra Stella

What Juno Includes

  • Create — prompt + lyrics (Write / Prompt / Instrumental) + style chips
    • More Options (Vocal Gender, Weirdness, Style Influence, Exclude), saved into workspaces, with live task status per song row.
  • Library — 11 tabs: Songs, Playlists, Workspaces, Studio Projects, Voices, Lyrics, Styles, Cover Art, Hooks, Liked Hooks, History; plus uploads (MP3/WAV/M4A/OGG/FLAC) and Trash with restore / delete forever.
  • Studio — multi-track arrangement view with region-based Generate / Repaint / Extend actions on the ACE-Step base model.
  • Editor — per-song waveform editing: Crop, Remove Section, Replace Section (ACE-Step repaint), Adjust Speed, Reverse, Export.
  • MIDI — audio → MIDI with MuScriptor (multi-instrument transcription by Kyutai × Mirelo). Upload audio or use ⋯ → Extract MIDI on any song; play it on a falling-notes piano (optionally over the original recording) or edit it in a piano roll, then download the .mid or render it back to audio.
  • Managed models — Juno keeps one XL model in VRAM, swaps it automatically per song, shows live status everywhere, and can free VRAM after an idle timeout. No "initialize" step, no refreshing.
  • Persistent player — full transport, queue, like/dislike, volume, track info.
  • Local proxy API on port 3000 that translates Juno requests into ACE-Step API calls (port 8001) and manages the on-disk library.
  • No accounts, credits, plans, ads, or telemetry. The profile is always "local user / Offline Mode".

Hardware Assumptions

  • NVIDIA GPU with 32 GB VRAM (the three exposed presets are all XL-class with the 4B language model; smaller ACE-Step variants exist but are intentionally not exposed in the UI).
  • ~60 GB free disk for model weights (downloaded on first run) plus space for outputs.
  • Docker with the NVIDIA Container Toolkit installed.
  • Linux host recommended. CUDA 12.4 base image.

Model Stack

Weights are not included in this ZIP and are not baked into the Docker image. On first start they are downloaded from Hugging Face into the host-mounted ./models directory:

Purpose Hugging Face repo Local path
ACE-Step 1.5 XL SFT (default) ACE-Step/acestep-v15-xl-sft /models/acestep-v15-xl-sft
ACE-Step 1.5 XL Turbo (8 steps, no CFG) ACE-Step/acestep-v15-xl-turbo /models/acestep-v15-xl-turbo
ACE-Step 1.5 XL Base (Studio tasks) ACE-Step/acestep-v15-xl-base /models/acestep-v15-xl-base
Language model (all presets) ACE-Step/acestep-5Hz-lm-4B /models/acestep-5Hz-lm-4B
MIDI transcription MuScriptor/muscriptor-{small,medium,large} (gated, CC BY-NC 4.0) HF cache, on first use
YuE2 generation + covers m-a-p/YuE2-3B, m-a-p/YuE2-Vae, m-a-p/YuE2-Vae-legacy /models/yue2/… (download yourself — see docs/YUE2.md)
Cover transcription m-a-p/SheetSage2 + m-a-p/MERT-v2-FullSong /models/yue2/…

Only one XL DiT is resident at a time (~9 GB each); switching presets swaps it. MuScriptor's weights are gated: accept the licence on the model page with the account that owns HF_TOKEN before your first transcription.

Quick Start

cd juno
cp .env.example .env      # edit paths / HF_TOKEN if needed
docker compose up --build

Then open http://localhost:3000.

To rebuild and recreate the container after updating Juno:

docker compose up -d --build --force-recreate juno
docker compose logs -f juno

If your Docker installation doesn't support docker compose GPU reservations, the equivalent manual run is:

docker build -t juno:latest .
docker run --gpus all \
  -p 3000:3000 -p 8001:8001 \
  -v "$(pwd)/models:/models" \
  -v "$(pwd)/outputs:/outputs" \
  -v "$(pwd)/uploads:/uploads" \
  -v "$(pwd)/data:/data" \
  -v "$(pwd)/hf-cache:/root/.cache/huggingface" \
  -e HF_TOKEN="${HF_TOKEN:-}" \
  juno:latest

First Run Behavior

  1. scripts/entrypoint.sh creates the storage directories (/outputs/library, /outputs/tmp, /outputs/cache/*, /uploads, /data).
  2. scripts/download_models.py downloads any missing model repos into /models (resumable; already-populated directories are skipped). The first download is tens of gigabytes — expect it to take a while.
  3. scripts/verify_models.py confirms all four model directories exist and are non-empty, then symlinks them into /app/ACE-Step-1.5/checkpoints/.
  4. supervisord starts the ACE-Step API (port 8001) and the Juno web/proxy server (port 3000).
  5. The UI is usable immediately. Models load on demand: the first Create loads the selected preset (about a minute), and the status card in the sidebar / Create panel shows every step. MuScriptor (supervisord program muscriptor, autostart off) starts on the first transcription and stops again after the idle timeout set in Settings.

Environment Variables

Set in .env (see .env.example) or the shell:

Variable Default Meaning
JUNO_MODEL_DIR ./models Host directory mounted at /models
JUNO_OUTPUT_DIR ./outputs Host directory mounted at /outputs
JUNO_UPLOAD_DIR ./uploads Host directory mounted at /uploads
JUNO_DATA_DIR ./data Host directory mounted at /data (library DB)
HF_HOME ./hf-cache Hugging Face cache mount
HF_TOKEN (empty) Hugging Face token. Required for MuScriptor (gated weights)
JUNO_LM_BACKEND pt 5Hz LM backend sent to ACE-Step (vllm is broken on Blackwell without flash-attn)
MUSCRIPTOR_MODEL medium Initial MIDI model size; switchable in the UI
JUNO_THINKING true false bypasses the 5Hz LM (pure DiT) for A/B debugging

Container-side variables (already set in docker-compose.yml): ACESTEP_API_HOST/PORT, ACESTEP_CONFIG_PATH (startup default model name; CONFIG_PATH2/3 are ignored — Juno swaps a single slot), ACESTEP_INIT_LLM=true, ACESTEP_LM_MODEL_PATH, ACESTEP_DEVICE=auto, ACESTEP_USE_FLASH_ATTENTION=true, ACESTEP_OFFLOAD_TO_CPU=false, ACESTEP_OFFLOAD_DIT_TO_CPU=false, ACESTEP_TMPDIR=/outputs/tmp, TRITON_CACHE_DIR, TORCHINDUCTOR_CACHE_DIR, and the JUNO_* mirrors.

Manual Model Download

If you prefer to pre-download (or the automatic download fails):

pip install "huggingface_hub>=0.34"   # >=0.34 provides the short `hf` CLI
export HF_TOKEN=...   # only if required

for repo in acestep-v15-xl-sft acestep-v15-xl-turbo acestep-v15-xl-base acestep-5Hz-lm-4B; do
  hf download "ACE-Step/${repo}" --local-dir "./models/${repo}"
done

(Older CLI versions: huggingface-cli download ... with the same arguments.) The container will detect the populated directories and skip downloading.

MuScriptor (MIDI transcription)

These weights are gated: accept the licence at https://huggingface.co/MuScriptor/muscriptor-<size> with the account that owns HF_TOKEN first, otherwise the download fails with 401/403.

Unlike the ACE-Step models, MuScriptor resolves its weights through the HuggingFace cache, so pre-download into the hub/ subdirectory of the cache you mount at /root/.cache/huggingface — not into ./models.

Use HF_HUB_CACHE, not HF_HOME: HF_HOME also relocates the token file ($HF_HOME/token), so overriding it makes an otherwise logged-in CLI anonymous and a gated repo then fails with 401 "Access denied".

# medium is the default; small / large are optional
HF_HUB_CACHE=./hf-cache/hub hf download MuScriptor/muscriptor-medium model.safetensors
HF_HUB_CACHE=./hf-cache/hub hf download MuScriptor/muscriptor-small  model.safetensors
HF_HUB_CACHE=./hf-cache/hub hf download MuScriptor/muscriptor-large  model.safetensors

This uses your hf auth login credentials. If you are not logged in, pass a token explicitly instead: HF_TOKEN=hf_... HF_HUB_CACHE=... hf download ….

The cache directory is created by the container as root, so a host-side download needs it to be writable by you first:

sudo chown -R "$(id -u):$(id -g)" ./hf-cache     # root inside the container is unaffected

Simplest alternative — run it inside the container, where the mount, the token and the permissions are all already correct:

docker compose exec juno hf download MuScriptor/muscriptor-medium model.safetensors

Otherwise the model downloads by itself on your first transcription, which is why that first run takes a while. Beat detection additionally fetches a small beat_this checkpoint from cloud.cp.jku.at into /models/torch-cache on first use; if that host is unreachable the transcription still succeeds, you just don't get the beat-quantized .mid.

Model Verification

docker compose exec juno python3 /app/juno/scripts/verify_models.py

Prints one line per model directory. All four must be present and non-empty: /models/acestep-v15-xl-sft, /models/acestep-v15-xl-turbo, /models/acestep-v15-xl-base, /models/acestep-5Hz-lm-4B.

Health Checks

# Juno proxy + aggregated ACE-Step status
curl http://localhost:3000/api/health

# Model preset status (available on disk / loaded in memory)
curl http://localhost:3000/api/models

# ACE-Step API directly
curl http://localhost:8001/health

/api/health reports {"juno":"ok","aceStep":"ok", ...} when everything is up. aceStep":"unavailable" while models are still downloading or the server is starting is expected.

Docker GPU Troubleshooting

Verify the NVIDIA runtime works at all:

docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
  • If that fails: install/repair the NVIDIA Container Toolkit (nvidia-ctk runtime configure --runtime=docker && sudo systemctl restart docker).
  • could not select device driver "nvidia" → toolkit not registered with Docker.
  • Driver/CUDA mismatch → host driver must support CUDA 12.4 (driver ≥ 550.x).
  • Inside the container: docker compose exec juno nvidia-smi.

HF Download Troubleshooting

  • 401/403 — set HF_TOKEN in .env (create one at huggingface.co/settings/tokens) and accept the model licenses on the repo pages if prompted.
  • 429 rate limit — wait and re-run; downloads resume where they left off.
  • Disk full — the four repos total tens of GB; check df -h ./models.
  • Corporate proxy/TLS issues — pre-download on another machine (see Manual Model Download) and copy the models/ folder over.
  • Partial downloads are safe: snapshot_download resumes, and non-empty completed directories are skipped.

Port Conflicts

Ports 3000 (web) and 8001 (ACE-Step) must be free. Change the host side of the mappings in docker-compose.yml (e.g. "3300:3000") if something else owns them; the in-container ports stay as-is.

Output Locations

Content Host path
Generated songs (library copies) ./outputs/library/
Export manifests ./outputs/exports/
MIDI transcriptions & edits ./outputs/midi/
Audio uploaded in the MIDI tab ./uploads/midi-src/
MuScriptor log ./outputs/cache/muscriptor.log
Uploaded audio ./uploads/
Library database ./data/juno-db.json
Model weights ./models/
HF cache ./hf-cache/
Temp / compile caches ./outputs/tmp, ./outputs/cache/

Upgrading

Cover behaviour, its controls and the research behind them: docs/COVER.md.

See UPGRADE.md for what changed in the MIDI / model-loading release and the one-time compose + licence steps.

Legal Note

Juno is an independent, offline project for personal/local use. It is not affiliated with Suno. ACE-Step models are downloaded directly from their Hugging Face repositories under their respective licenses — review those licenses (including any restrictions on generated-output usage) before distributing anything you create. MuScriptor weights and their output are CC BY-NC 4.0 (non-commercial) and require that you have the rights to any audio you transcribe. You are responsible for the content you generate with your own hardware.

Running Without Models (UI-only mode)

To explore the full UI without downloading ~60 GB of weights (or on a machine without a suitable GPU):

JUNO_SKIP_MODELS=true docker compose up --build

Everything renders and works from mock data — pages, player, library tabs, editor, studio, all song-row states. ACE-Step shows as "unavailable", the ACE-Step process may restart-loop harmlessly in the logs, and any generation attempt creates a documented failed row instead of audio. Later, run normally (without the flag) and the models download; nothing else changes. Local-only actions (Crop, Reverse, Speed, uploads, exports, playlists) work fully in this mode.

About

A suno.com clone (work-in-progress)

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages