KoshiMazaki commited on
Commit
34b771b
·
verified ·
1 Parent(s): 4a7b166

AKUSPACE LTX-2.5 Audio LoRA v0.5 — checkpoint 11500, card, config, six A/B examples

Browse files
.gitattributes CHANGED
@@ -33,3 +33,11 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ examples/beat-dual-delay.mp3 filter=lfs diff=lfs merge=lfs -text
37
+ examples/beat-empty-club.mp3 filter=lfs diff=lfs merge=lfs -text
38
+ examples/beat-medium-room.mp3 filter=lfs diff=lfs merge=lfs -text
39
+ examples/beat-outdoor-day.mp3 filter=lfs diff=lfs merge=lfs -text
40
+ examples/dry-beat.mp3 filter=lfs diff=lfs merge=lfs -text
41
+ examples/dry-voice.mp3 filter=lfs diff=lfs merge=lfs -text
42
+ examples/voice-cathedral.mp3 filter=lfs diff=lfs merge=lfs -text
43
+ examples/voice-small-room.mp3 filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,199 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: ltx-2-community-license-agreement
4
+ license_link: https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md
5
+ base_model: Lightricks/LTX-2.5
6
+ pipeline_tag: audio-to-audio
7
+ tags:
8
+ - ltx
9
+ - ltx-2.5
10
+ - lora
11
+ - audio-to-audio
12
+ - audio
13
+ - spatial-audio
14
+ - room-acoustics
15
+ - reverb
16
+ - ambience
17
+ - comfyui
18
+ ---
19
+
20
+ # AKUSPACE — LTX-2.5 Audio LoRA v0.5
21
+
22
+ **Acoustic space control for LTX-2.5.** Give it a dry recording and a prompt naming
23
+ an acoustic treatment, and it re-renders that recording as though it were captured
24
+ in that space — while the performance stays put on the timeline.
25
+
26
+ Trigger word: `AKUSPACE`
27
+
28
+ - Website and A/B examples: https://akuspace.pages.dev
29
+ - ComfyUI nodes: https://github.com/koshimazaki/ComfyUI-Koshi-Nodes
30
+
31
+ ## What it does
32
+
33
+ Three kinds of treatment, each with a text-controlled amount:
34
+
35
+ | Mode | What it is | Options |
36
+ |---|---|---|
37
+ | **Space** | rooms a sound sits *in* | small room (0.67 s), medium room (1.9 s), empty club (1.2 s), cathedral |
38
+ | **Place** | environments a sound sits *among* | outdoor day, outdoor night |
39
+ | **Sound effects** | processing a sound goes *through* | modular granular delay |
40
+
41
+ Each treatment has a level axis, and the level word is part of the trained caption:
42
+
43
+ - **gentle** — subtle inclusion of the effect
44
+ - **moderate** — the sweet spot, and the level to showcase
45
+ - **heavy** — pushes the dry signal into sound-effect territory. Shipped as
46
+ **experimental**: it is the edge of the trained range, not a general-purpose setting.
47
+
48
+ Rooms and sound effects carry all three amounts. **Outdoor places carry only
49
+ `gentle` and `heavy`** — an ambience bed is a separate recording rather than a
50
+ tail, so it survives being scaled down but not up.
51
+
52
+ The level word occupies the same position in every caption, which is what makes it
53
+ a consistent semantic control rather than an arbitrary number.
54
+
55
+ ## It holds sync with picture
56
+
57
+ Envelope cross-correlation of each published example against its own dry reference,
58
+ measured on this checkpoint (step 11500), 2 ms resolution:
59
+
60
+ | Example | Offset | Peak correlation |
61
+ |---|---|---|
62
+ | Voice · small room | −20 ms | 0.96 |
63
+ | Voice · cathedral | −22 ms | 0.97 |
64
+ | Beat · medium room | −16 ms | 0.34 |
65
+ | Beat · empty club | −20 ms | 0.30 |
66
+ | Beat · outdoor day | −24 ms | 0.32 |
67
+ | Beat · dual delay | −22 ms | 0.34 |
68
+
69
+ The offset is **16–24 ms across every mode and both source types** — a consistent,
70
+ small lead rather than per-clip drift, and inside lip-sync tolerance. So a
71
+ performance can be moved from a studio into a cathedral in post, by describing it,
72
+ without re-shooting or re-syncing.
73
+
74
+ **Read the correlation column carefully.** It measures how much the amplitude
75
+ envelope *changed*, and reverb changes it by design — most of all on
76
+ transient-dense material, where the tail fills the gaps between hits. The low
77
+ figures on percussion are the expected signature of the effect working, not a
78
+ defect, and they are reported here rather than omitted. Envelope correlation is a
79
+ timing measure, not a quality measure; judge quality by ear.
80
+
81
+ ## Use
82
+
83
+ ```text
84
+ AKUSPACE female spoken voice through synthetic cathedral reverb, moderate
85
+ ```
86
+
87
+ Supply the corresponding dry audio as the **audio-to-audio reference**.
88
+
89
+ Suggested inference settings: **24 steps, CFG 4**.
90
+
91
+ > **Turn prompt enhancement OFF.** The rewriter paraphrases away the trigger word
92
+ > and the level word, which are exactly the tokens the control surface depends on.
93
+
94
+ **This adapter is reference-conditioned.** It was trained only to transform a
95
+ supplied reference, so plain text-to-audio with the LoRA loaded and no reference
96
+ present generates near-silence at any strength. That is a usage constraint, not a
97
+ defect. For video, the working chain is: dry voice → audio-to-audio space pass →
98
+ image+audio-to-video, with the audio held fixed and the base model doing lip-sync.
99
+
100
+ ## Training
101
+
102
+ - Base checkpoint: `Lightricks/LTX-2.5` (22B dev transformer, bf16)
103
+ - LoRA rank / alpha: 32 / 32, dropout 0.0
104
+ - Target modules: `audio_attn1`, `audio_attn2`, `audio_ff` — the audio branches
105
+ only. The video branch and both cross-modal attention paths are untouched, which
106
+ is why the spatial effect survives inside joint audio-video renders.
107
+ - Optimiser: AdamW, learning rate 2e-4, batch size 1, bf16, seed 42
108
+ - Steps: trained to 12,300; **this release ships step 11,500**, chosen by listening
109
+ across a 228-clip sweep of four candidate checkpoints against every trained cell.
110
+ - Dataset: 266 paired clips — 228 train / 38 validation
111
+ - Clip length 6.000 s, 48 kHz, −3 dBFS ceiling
112
+ - Held out from training: one beat and one male voice source
113
+
114
+ ## How the dataset was built
115
+
116
+ Every item is **the same performance twice** — once dry, once through a real
117
+ acoustic treatment — so the model learns a transformation rather than an
118
+ association.
119
+
120
+ The material is my own, recorded and produced over several years: multiple speaking
121
+ voices, beats and electronic music, and acoustic instruments (piano, saxophone,
122
+ citola, handclaps). Treatments come from digital reverbs, analog and modular
123
+ patches, and original field recordings for the outdoor beds — captured as spaces,
124
+ as places, as sound effects, and across different room sizes. Nothing is scraped or
125
+ third-party licensed.
126
+
127
+ Sources were rendered as a deliberate grid — 14 sources × (4 rooms × 3 levels + one
128
+ effect chain × 3 levels + 2 places × 2 levels) — so that no space is defined by one
129
+ voice and no voice is defined by one space. Captions name the space and its decay
130
+ time, putting the physical quantity into the text, which is what makes a continuous
131
+ control surface reachable.
132
+
133
+ Two decisions did more for quality than any hyperparameter:
134
+
135
+ **Source coverage per space.** An earlier version trained on 8 sources produced
136
+ reverb that read as a lo-fi proxy. Raising per-space coverage to 14, with no other
137
+ change, is what made the rooms believable. Coverage, not capacity.
138
+
139
+ **Not fading the tails.** Clips are cut at 6 seconds without a fade, leaving the
140
+ reverb tail faintly ringing at the boundary (−44 to −76 dB). Fading to silence
141
+ teaches the model that reverb *stops* at the window edge; leaving it ringing
142
+ teaches that it continues, which is what lets a 6-second-trained adapter generate a
143
+ coherent 15-second tail.
144
+
145
+ A verification script gates the pipeline before every training run: grid
146
+ completeness, byte-level file integrity, sample-rate and channel consistency,
147
+ per-source level spread, and per-space processing depth. Each of those checks
148
+ caught a real failure. The integrity check compares declared frame counts against
149
+ audio data actually present — a truncated WAV keeps its original header, so every
150
+ duration-based check reports a healthy file while the encoder silently skips it.
151
+
152
+ ## Limitations
153
+
154
+ - **Sung vocals are out of distribution.** Every trained voice is *spoken*.
155
+ - **Brass is out of distribution.** Trumpet appears only inside a single combined
156
+ piano-and-saxophone source. Results on solo brass are generalisation, not a
157
+ trained capability.
158
+ - Requires an audio reference; it will not generate a space from text alone.
159
+ - This is generative transformation, not physically accurate acoustic simulation.
160
+ - It may alter wording, timing, pitch, or timbre — audio is regenerated, not
161
+ filtered.
162
+ - Outdoor ambience can mask a quiet source, and beds are recordings rather than
163
+ reverb tails, so they scale down more gracefully than up.
164
+ - The level axis is a learned caption control, not a calibrated wet/dry percentage.
165
+
166
+ ## Files
167
+
168
+ | Path | What it is |
169
+ |---|---|
170
+ | `akuspace-ltx25-v0.5.safetensors` | The adapter — step 11500, bf16, rank 32 |
171
+ | `config/a2a_v5_ltx25.yaml` | The training config this run used |
172
+ | `examples/` | The six A/B pairs, 320k MP3 |
173
+
174
+ The examples share two dry references. `dry-voice.mp3` pairs with
175
+ `voice-small-room.mp3` and `voice-cathedral.mp3`; `dry-beat.mp3` pairs with
176
+ `beat-medium-room.mp3`, `beat-empty-club.mp3`, `beat-outdoor-day.mp3` and
177
+ `beat-dual-delay.mp3`. Rooms are at `moderate`, outdoor at `gentle`, dual delay at
178
+ `heavy`. No gain or normalisation was applied on export, so the level relationship
179
+ between dry and processed is the one the model produced.
180
+
181
+ An interactive crossfade between each pair is at https://akuspace.pages.dev
182
+
183
+ ## Examples and provenance
184
+
185
+ The published examples use a synthetic TTS voice and an AI-generated beat as dry
186
+ sources. That is deliberate rather than a compromise: a replacement or synthetic
187
+ voice track that needs to be seated in a scene is the realistic use case for this
188
+ adapter, and it is the case the audio-dub category is about.
189
+
190
+ Training data contains no synthetic-TTS-vendor material; the voices in the training
191
+ grid are my own recordings.
192
+
193
+ ## Licence
194
+
195
+ Released under the **LTX-2.x Community License** terms that govern the base model.
196
+ Confirm the current terms at
197
+ https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md before commercial use —
198
+ the community licence covers entities under $10M annual revenue, with separate
199
+ commercial agreements above that threshold.
akuspace-ltx25-v0.5.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:623a32970c0984a80af5364d5257d1d3a7fa0c8fc112caee427d7aa631b3e361
3
+ size 163714520
config/a2a_v5_ltx25.yaml ADDED
@@ -0,0 +1,112 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ model:
2
+ model_path: /workspace/models/ltx-2.5/diffusion_models/ltx-2.5-22b-dev-transformer-bf16.safetensors
3
+ text_encoder_path: /workspace/models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
4
+ training_mode: lora
5
+ load_checkpoint: null
6
+ video_vae_path: /workspace/models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors
7
+ audio_vae_path: /workspace/models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors
8
+ lora:
9
+ rank: 32
10
+ alpha: 32
11
+ dropout: 0.0
12
+ target_modules:
13
+ - audio_attn1.to_k
14
+ - audio_attn1.to_q
15
+ - audio_attn1.to_v
16
+ - audio_attn1.to_out.0
17
+ - audio_attn2.to_k
18
+ - audio_attn2.to_q
19
+ - audio_attn2.to_v
20
+ - audio_attn2.to_out.0
21
+ - audio_ff.net.0.proj
22
+ - audio_ff.net.2
23
+ training_strategy:
24
+ name: flexible
25
+ audio:
26
+ is_generated: true
27
+ latents_dir: audio_latents
28
+ conditions:
29
+ - type: reference
30
+ latents_dir: reference_audio_latents
31
+ probability: 1.0
32
+ optimization:
33
+ learning_rate: 0.0002
34
+ steps: 12300
35
+ batch_size: 1
36
+ gradient_accumulation_steps: 1
37
+ max_grad_norm: 1.0
38
+ optimizer_type: adamw
39
+ scheduler_type: linear
40
+ scheduler_params: {}
41
+ enable_gradient_checkpointing: true
42
+ acceleration:
43
+ mixed_precision_mode: bf16
44
+ quantization: null
45
+ load_text_encoder_in_8bit: false
46
+ offload_optimizer_during_validation: false
47
+ data:
48
+ preprocessed_data_root: /workspace/Demos/data/acoustic-space-v5/.precomputed
49
+ num_dataloader_workers: 2
50
+ validation:
51
+ samples:
52
+ - prompt: AKUSPACE male spoken voice in a small bathroom-like room, moderate bright close reflections and a short 0.67-second reverb decay, no background ambience
53
+ conditions:
54
+ - type: reference
55
+ audio: /workspace/Demos/data/acoustic-space-v5/audio/references/male_voice_02_small_room_0_67_mid.wav
56
+ - prompt: AKUSPACE male spoken voice through synthetic cathedral reverb, moderate wide diffuse reflections and a long decaying tail, no background ambience
57
+ conditions:
58
+ - type: reference
59
+ audio: /workspace/Demos/data/acoustic-space-v5/audio/references/male_voice_02_synthetic_cathedral_mid.wav
60
+ - prompt: AKUSPACE electronic rhythm loop outdoors in daytime, heavy open-air acoustics with continuous birdsong ambience
61
+ conditions:
62
+ - type: reference
63
+ audio: /workspace/Demos/data/acoustic-space-v5/audio/references/beat_02_outdoor_day_birds_high.wav
64
+ - prompt: AKUSPACE electronic rhythm loop outdoors at night, heavy open-air acoustics with crickets and distant car ambience
65
+ conditions:
66
+ - type: reference
67
+ audio: /workspace/Demos/data/acoustic-space-v5/audio/references/beat_02_outdoor_night_high.wav
68
+ - prompt: AKUSPACE male spoken voice through a modular granular delay, moderate scattered grains and unpredictable modulated echoes, no background ambience
69
+ conditions:
70
+ - type: reference
71
+ audio: /workspace/Demos/data/acoustic-space-v5/audio/references/male_voice_02_modular_granular_delay_mid.wav
72
+ negative_prompt: distorted, clipped, metallic, unintelligible speech, changed words, changed rhythm
73
+ video_dims:
74
+ - 512
75
+ - 512
76
+ - 153
77
+ frame_rate: 25.0
78
+ seed: 42
79
+ inference_steps: 24
80
+ interval: 250
81
+ guidance_scale: 4.0
82
+ stg_scale: 1.0
83
+ stg_blocks:
84
+ - 29
85
+ stg_mode: stg_av
86
+ generate_audio: true
87
+ generate_video: false
88
+ skip_initial_validation: true
89
+ checkpoints:
90
+ interval: 100
91
+ keep_last_n: -1
92
+ precision: bfloat16
93
+ flow_matching:
94
+ timestep_sampling_mode: shifted_logit_normal
95
+ timestep_sampling_params: {}
96
+ hub:
97
+ push_to_hub: false
98
+ hub_model_id: null
99
+ wandb:
100
+ enabled: false
101
+ project: ltx-acoustic-space-control
102
+ entity: null
103
+ tags:
104
+ - ltx2.3
105
+ - a2a
106
+ - ic-lora
107
+ - audio
108
+ - acoustic-space
109
+ - ableton
110
+ log_validation_videos: true
111
+ seed: 42
112
+ output_dir: /workspace/Demos/runs/acoustic-space-v5-ltx25
examples/beat-dual-delay.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:de2a1a421a63edf1e20a1e973d6f5b769fce3cd934c5b075004f7bde39733110
3
+ size 604844
examples/beat-empty-club.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b8c74cf6a899cf51a1b7be9164c4c00048daf0d0694319163bf8bb4f7898dc78
3
+ size 604844
examples/beat-medium-room.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fd11d77cb89b59dafabee0adb1ad252191fc0a4832a986f0df6a49b0dc3940e0
3
+ size 604844
examples/beat-outdoor-day.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:34057000f8d7235b8c027303bc7d59592c834a9fdc23a028e00dae54f719b98f
3
+ size 604844
examples/dry-beat.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d08c33c6d978c4577b7189ede5c4d3881381afda7eca3e2adaa9bfcc383df2ca
3
+ size 605804
examples/dry-voice.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:36c1680c582f02ef8507278c0ae7f9aaf847cea97573a7be7cc450f937642f86
3
+ size 605804
examples/voice-cathedral.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4b1203f9b50f931e2c9e2b1c492dcb4ce3c499da0c4546c0bc0427146ff9e6f0
3
+ size 604844
examples/voice-small-room.mp3 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a8919b5fbcfe1ad55b34329b306ef5d58d51250ea66125cfcda22aff55b056d1
3
+ size 604844