droyster commited on
Commit
4d45bde
·
verified ·
1 Parent(s): 67196d0

Use upstream Meta Llama tokenizer default

Browse files
README.md CHANGED
@@ -29,7 +29,7 @@ This repository contains a 4-bit TorchAO weight-only quantization of
29
  - **Quantization:** TorchAO `Int4WeightOnlyConfig(group_size=128)`
30
  - **Runtime format:** `torch.save` checkpoint containing TorchAO quantized tensor subclasses
31
  - **Tested GPU:** RTX 3060 12 GB
32
- - **Tokenizer override used for testing:** `unsloth/Llama-3.2-1B`
33
  - **Language:** English, following the base model
34
 
35
  No private prompt voice is included. Voice continuation/cloning requires user-supplied prompt audio and transcript.
@@ -50,11 +50,7 @@ source .venv/bin/activate
50
  pip install -r requirements.txt
51
  ```
52
 
53
- If the upstream tokenizer is gated/unavailable in your environment, set:
54
-
55
- ```bash
56
- export MISO_TTS_TOKENIZER=unsloth/Llama-3.2-1B
57
- ```
58
 
59
  ## Quick smoke test
60
 
 
29
  - **Quantization:** TorchAO `Int4WeightOnlyConfig(group_size=128)`
30
  - **Runtime format:** `torch.save` checkpoint containing TorchAO quantized tensor subclasses
31
  - **Tested GPU:** RTX 3060 12 GB
32
+ - **Tokenizer:** upstream default `meta-llama/Llama-3.2-1B`
33
  - **Language:** English, following the base model
34
 
35
  No private prompt voice is included. Voice continuation/cloning requires user-supplied prompt audio and transcript.
 
50
  pip install -r requirements.txt
51
  ```
52
 
53
+ The loader uses upstream MisoTTS tokenizer behavior by default: `meta-llama/Llama-3.2-1B`. This requires a Hugging Face token/account that has access to Meta's Llama 3.2 tokenizer repo.
 
 
 
 
54
 
55
  ## Quick smoke test
56
 
__pycache__/generator.cpython-310.pyc ADDED
Binary file (7.54 kB). View file
 
__pycache__/load_quantized.cpython-310.pyc ADDED
Binary file (2.83 kB). View file
 
__pycache__/models.cpython-310.pyc ADDED
Binary file (8.68 kB). View file
 
__pycache__/moshi_compat.cpython-310.pyc ADDED
Binary file (1.57 kB). View file
 
__pycache__/watermarking.cpython-310.pyc ADDED
Binary file (2.72 kB). View file
 
config.json CHANGED
@@ -9,5 +9,5 @@
9
  "quantized_modules": "nn.Linear.weight",
10
  "runtime": "pytorch+torchao"
11
  },
12
- "tokenizer_override_tested": "unsloth/Llama-3.2-1B"
13
  }
 
9
  "quantized_modules": "nn.Linear.weight",
10
  "runtime": "pytorch+torchao"
11
  },
12
+ "tokenizer": "meta-llama/Llama-3.2-1B"
13
  }
scripts/__pycache__/export_int4.cpython-310.pyc ADDED
Binary file (4.77 kB). View file
 
scripts/__pycache__/smoke_test.cpython-310.pyc ADDED
Binary file (1.85 kB). View file
 
scripts/export_int4.py CHANGED
@@ -10,7 +10,7 @@ from pathlib import Path
10
  os.environ.setdefault("HF_HUB_ETAG_TIMEOUT", "60")
11
  os.environ.setdefault("HF_HUB_DOWNLOAD_TIMEOUT", "60")
12
  os.environ.setdefault("NO_TORCH_COMPILE", "1")
13
- os.environ.setdefault("MISO_TTS_TOKENIZER", "unsloth/Llama-3.2-1B")
14
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
15
 
16
  import torch
 
10
  os.environ.setdefault("HF_HUB_ETAG_TIMEOUT", "60")
11
  os.environ.setdefault("HF_HUB_DOWNLOAD_TIMEOUT", "60")
12
  os.environ.setdefault("NO_TORCH_COMPILE", "1")
13
+ os.environ.setdefault("MISO_TTS_TOKENIZER", "meta-llama/Llama-3.2-1B")
14
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
15
 
16
  import torch
scripts/smoke_test.py CHANGED
@@ -6,7 +6,7 @@ import os
6
  from pathlib import Path
7
 
8
  os.environ.setdefault("NO_TORCH_COMPILE", "1")
9
- os.environ.setdefault("MISO_TTS_TOKENIZER", "unsloth/Llama-3.2-1B")
10
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
11
 
12
  import torch
 
6
  from pathlib import Path
7
 
8
  os.environ.setdefault("NO_TORCH_COMPILE", "1")
9
+ os.environ.setdefault("MISO_TTS_TOKENIZER", "meta-llama/Llama-3.2-1B")
10
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
11
 
12
  import torch