UNDERSCORE
AI That Scores Your Video, Scene by Scene
動画をシーンごとに読み取り、音楽をつけるAI
Overview
概要Finding the right music for a video is slow and licence-fraught. Underscore watches your footage, understands each scene, and generates an original soundtrack matched to the cuts — composed locally, on your own machine.
動画に合う音楽を探すのは時間がかかり、ライセンスの問題もつきまといます。Underscoreは映像を見てシーンごとの内容を理解し、カットに合わせたオリジナルのサウンドトラックを生成します。作曲は自分のマシン上で、ローカルに行います。
An ambitious experiment: can a fully local AI pipeline actually compose to picture, the way a human composer scores to an edit?
挑戦的な実験です。人間の作曲家が編集に合わせて音楽をつけるように、完全にローカルなAIのパイプラインで映像に合わせた作曲ができるのかを試しました。
How it was built
開発It segments the video, uses Gemini to read the mood and content of each scene, then drives a local ACE-Step music model (GGUF) to generate audio — aligning everything to the edit with FFmpeg, with YAMNet assisting the analysis.
動画をシーンに分割し、Geminiで各シーンの雰囲気と内容を読み取ったうえで、ローカルの音楽生成モデルACE-Step(GGUF)で音声を生成します。FFmpegで編集に合わせて全体を整え、映像の解析にはYAMNetも使っています。
What it does
機能Gemini reads the mood and content of each scene, producing the brief the music generator works from.
Geminiが各シーンの雰囲気と内容を読み取り、音楽生成のもとになる指示を作ります。
A local ACE-Step diffusion model (GGUF) composes an original score on-device — no licensing, no cloud.
ローカルの拡散モデルACE-Step(GGUF)が、端末上でオリジナルの音楽を作曲します。ライセンスもクラウドも不要です。
Audio is aligned to the edit with FFmpeg, with YAMNet helping analyse the footage's own sound.
FFmpegで音声を編集に合わせ、YAMNetで映像に含まれる音の解析を補助します。
Tauri 2 + SvelteKit 5, Apple-Silicon-first, with the whole pipeline running locally.
Tauri 2+SvelteKit 5で、Apple Siliconを優先して対応。パイプライン全体がローカルで動きます。
Under the hood
技術の詳細Understanding feeds generation理解を生成につなげる
Underscore segments a video, uses Gemini to read the mood and content of each scene, and turns that into the brief a music model composes from. It's multimodal understanding driving generation, scene by scene.
Underscoreは動画をシーンに分割し、Geminiで各シーンの雰囲気と内容を読み取り、それを音楽モデルが作曲するための指示に変えます。シーンごとに、マルチモーダルな理解が生成を動かします。
Local generative musicローカルでの音楽生成
A local ACE-Step diffusion model (GGUF) generates an original score on-device — no licensing, no cloud — with FFmpeg aligning audio to the cuts and YAMNet assisting analysis. An early demo, Apple-Silicon-first, and exactly the kind of pipeline I find worth chasing.
ローカルの拡散モデルACE-Step(GGUF)が端末上でオリジナルの音楽を生成します。ライセンスもクラウドも不要で、FFmpegが音声をカットに合わせ、YAMNetが解析を補助します。Apple Siliconを優先した初期段階のデモで、まさに追いかける価値があると思える種類のパイプラインです。
Stack
技術構成Underscore is Tauri 2 + SvelteKit 5 + Rust. Gemini handles per-scene mood analysis; a local ACE-Step diffusion music model (GGUF) generates audio on-device; FFmpeg and YAMNet handle alignment and audio analysis. It is an early demo, Apple-Silicon-only today.
UnderscoreはTauri 2+SvelteKit 5+Rustです。シーンごとの雰囲気の解析はGemini、音声の生成はローカルで動く拡散モデルACE-Step(GGUF)、位置合わせと音声解析はFFmpegとYAMNetが担います。初期段階のデモで、現在はApple Siliconのみに対応しています。
Where it is now
現状と今後Underscore is an early, ambitious demo: a working Gemini-plus-ACE-Step pipeline that scores video on-device, Apple-Silicon-first. Longer clips and finer control over the generated score are where it goes next — the multimodal-to-generative loop is the hard part, and it already runs.
Underscoreは初期段階の挑戦的なデモです。GeminiとACE-Stepを組み合わせたパイプラインが、Apple Siliconを優先して端末上で動画に音楽をつけるところまで動いています。次は、より長い動画への対応と、生成する音楽の細かな調整です。難しいのはマルチモーダルな理解から生成へとつなぐ部分で、それはすでに動いています。
Underscore is the most ambitious pipeline here — multimodal understanding feeding a local generative model. Early and rough, and exactly the kind of problem I find worth chasing.
Underscoreは、この中でいちばん挑戦的なパイプラインです。マルチモーダルな理解をローカルの生成モデルにつなげています。まだ初期段階で粗削りですが、まさに追いかける価値があると思える種類の課題です。