nimafathi commited on
Commit
f1cceec
·
verified ·
1 Parent(s): fc3f5cb

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +105 -96
README.md CHANGED
@@ -2,123 +2,132 @@
2
  language:
3
  - en
4
  tags:
 
5
  - diffusion-language-model
6
- - dllms
7
  - text-generation
8
  - diffusion
9
  - language-model
10
- license: mit
11
  ---
12
 
13
- # hdlm-group/hdlm-base-epsilon-0.05
14
-
15
- This is a epsilon_hybrid diffusion language model trained on text data.
16
-
17
- ## Model Details
18
-
19
- - **Model Type**: epsilon_hybrid
20
- - **Architecture**: Diffusion-based language model
21
- - **Training Method**: Epsilon-hybrid diffusion training
22
-
23
- ## Configuration
24
-
25
- ```yaml
26
- hf_model_id: hdlm-group/hdlm-base-epsilon-0.0
27
- reset_step_for_finetuning: true
28
- ngpus: 4
29
- type: aligned
30
- gradient_accumulation_steps: 8
31
- model_type: epsilon_hybrid
32
- tokenizer:
33
- tokens: 50257
34
- model: gpt2
35
- training:
36
- batch_size: 512
37
- accum: ${gradient_accumulation_steps}
38
- n_iters: 500000
39
- snapshot_freq: 5000
40
- log_freq: 500
41
- eval_freq: 5000
42
- snapshot_freq_for_preemption: 1000
43
- snapshot_sampling: true
44
- ema: 0.9999
45
- warmup_iter: 50000
46
- loss_type: hybrid
47
- epsilon: 0.05
48
- lambda: 5.0
49
- lr: 1.0e-05
50
- data:
51
- train: openwebtext-train
52
- valid: wikitext103
53
- cache_dir: /home/toolkit/research-diffcodegen/data
54
- debug: false
55
- annealing:
56
- type: none
57
- efficient: false
58
- width: 1024
59
- tau: 1024
60
- eval_tau: 1024
61
- sampling_method: sdlm
62
- sampling_eps: 0.0001
63
- attention:
64
- context_type: block_causal
65
- block_type: full
66
- match_inference: true
67
- eval:
68
- batch_size: 32
69
- perplexity: true
70
- perplexity_batch_size: 16
71
- optim:
72
- weight_decay: 0.1
73
- optimizer: AdamW
74
- lr: 5.0e-05
75
- beta1: 0.9
76
- beta2: 0.95
77
- eps: 1.0e-08
78
- warmup: 10000
79
- grad_clip: 1.0
80
- scheduler: cosine
81
- experiment:
82
- name: ft_epsilon_0.05_lambda_5.0
83
- wandb_project: Hybrid-SDLM-ALIGNED
84
- model:
85
- name: epsilon_hdlm
86
- type: ddit
87
- hidden_size: 768
88
- cond_dim: 128
89
- length: 1024
90
- n_blocks: 12
91
- n_heads: 12
92
- dropout: 0.1
93
- scale_by_sigma: false
94
- transformer_sigma_conditioning: false
95
- hybrid_sigma_embedding: false
96
- post_process_logits: false
97
- use_timestep_embedding: false
98
 
99
- ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
100
 
101
  ## Usage
102
 
 
 
103
  ```python
104
- from our.hf_utils import smart_model_loader
 
 
 
 
 
 
 
 
 
 
105
 
106
- # Load the model
107
- model, config, device, accelerator, metaschedule = smart_model_loader(
108
- "hdlm-group/hdlm-base-epsilon-0.05",
109
- model_type="epsilon_hybrid"
 
 
 
 
 
 
 
 
 
 
 
 
 
110
  )
111
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
112
  ```
113
 
114
  ## Training Details
115
 
116
- please refer to the official GitHub Repository: https://github.com/ServiceNow/hdlm
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
117
 
118
  ## Citation
119
 
120
- If you use this model in your research, please cite the original paper and this implementation.
 
 
 
 
 
 
 
121
 
122
  ## License
123
 
124
- This model is released under the Apache License Version 2.0.
 
2
  language:
3
  - en
4
  tags:
5
+ - dllm
6
  - diffusion-language-model
 
7
  - text-generation
8
  - diffusion
9
  - language-model
10
+ license: apache-2.0
11
  ---
12
 
13
+ # HDLM-Epsilon: Hybrid Diffusion Language Model
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
+ [![Paper](https://img.shields.io/badge/Paper-arXiv-red)](https://arxiv.org/abs/2504.06416)
16
+ [![Code](https://img.shields.io/badge/Code-GitHub-blue)](https://github.com/ServiceNow/hdlm)
17
+
18
+ This model card is for the **hdlm-base model with epsilon=0.05**
19
+
20
+ ## Model Description
21
+
22
+ HDLM-Epsilon is a hybrid diffusion language model that unifies autoregressive and diffusion-based sequence generation through epsilon-hybrid noising. This model interpolates evolution operators between absorbing and uniform processes, making it conceptually closer to MDLM (Sahoo et al. 2024) while maintaining the benefits of both paradigms.
23
+
24
+ The epsilon parameter (ε) controls the blend between absorbing and uniform processes during training, where smaller values emphasize the absorbing process and larger values incorporate more uniform noise.
25
+
26
+ ## Model Architecture
27
+
28
+ - **Base Model**: Transformer architecture with custom conditioning layers
29
+ - **Vocabulary Size**: 50,258 tokens (GPT-2 vocabulary + absorbing token)
30
+ - **Context Length**: 1024 tokens
31
+ - **Training**: Hybrid loss combining token masking with random token corruption
32
+ - **Inference**: Supports multiple sampling algorithms including ACS (Adaptive Correction Sampler)
33
 
34
  ## Usage
35
 
36
+ ### Quick Start
37
+
38
  ```python
39
+ from hdlm.hf_utils import smart_model_loader
40
+ from hdlm.epsilon_hybrid.sample import full_diff
41
+ from transformers import GPT2TokenizerFast
42
+ import torch
43
+
44
+ # Load model using smart loader (automatically detects model type)
45
+ model, cfg, device, accelerator, metaschedule = smart_model_loader(
46
+ model_path="hdlm-group/hdlm-base-epsilon-0.05",
47
+ model_type="auto", # automatically detects epsilon_hybrid
48
+ device="cuda"
49
+ )
50
 
51
+ # Load tokenizer
52
+ tokenizer = GPT2TokenizerFast.from_pretrained('gpt2')
53
+
54
+ # Generate text
55
+ prompt = "The future of artificial intelligence"
56
+ prompt_ids = tokenizer.encode(prompt, return_tensors='pt').to(device)
57
+
58
+ # Full diffusion sampling
59
+ generated = full_diff(
60
+ model=model,
61
+ prompt=prompt_ids,
62
+ batch_size=1,
63
+ alg='acs', # or 'original', 'remask', 'remdm'
64
+ steps=512,
65
+ temperature=1.0,
66
+ context_length=1024,
67
+ device=device
68
  )
69
 
70
+ # Decode generated text
71
+ generated_text = tokenizer.decode(generated[0], skip_special_tokens=True)
72
+ print(generated_text)
73
+ ```
74
+
75
+ ### Evaluation
76
+
77
+ ```bash
78
+ # Text generation evaluation
79
+ python hdlm/eval_generation.py \
80
+ --checkpoint_path hdlm-group/hdlm-base-epsilon-0.05 \
81
+ --sampling_method full_diff \
82
+ --algorithm acs \
83
+ --save_samples
84
+
85
+ # Perplexity evaluation
86
+ python hdlm/eval_modeling.py \
87
+ --checkpoint_path hdlm-group/hdlm-base-epsilon-0.05 \
88
+ --work_dir "./logs/eval_modeling_epsilon" \
89
+ --dataset ptb
90
  ```
91
 
92
  ## Training Details
93
 
94
+ - **Dataset**: OpenWebText
95
+ - **Batch Size**: 512
96
+ - **Learning Rate**: 3e-4 with cosine scheduling
97
+ - **Epsilon (ε)**: 0.01 (controls hybrid noising blend)
98
+ - **Lambda (λ)**: 1.0 (weighting factor for unmasked tokens)
99
+ - **Loss Type**: Hybrid loss combining masking and random token corruption
100
+ - **Training Steps**: 1M iterations
101
+ - **Warmup**: 50K steps
102
+
103
+ ## Sampling Algorithms
104
+
105
+ The model supports several sampling algorithms:
106
+
107
+ - **`original`**: Standard diffusion sampling
108
+ - **`acs`**: Adaptive Correction Sampler with error correction
109
+ - **`remask`**: Remasking strategy for improved quality
110
+ - **`remdm`**: ReMDM-style sampling with probability mixing
111
+
112
+ ## Model Variants
113
+
114
+ Available epsilon values and their characteristics:
115
+
116
+ - **ε = 0.01**: Minimal uniform noise, closest to pure absorbing process
117
+ - **ε = 0.1**: Moderate hybrid behavior
118
+ - **ε = 0.5**: Balanced absorbing-uniform blend
119
 
120
  ## Citation
121
 
122
+ ```bibtex
123
+ @article{fathi2025unifying,
124
+ title={Unifying autoregressive and diffusion-based sequence generation},
125
+ author={Fathi, Nima and Scholak, Torsten and No{\"e}l, Pierre-Andr{\'e}},
126
+ journal={arXiv preprint arXiv:2504.06416},
127
+ year={2025}
128
+ }
129
+ ```
130
 
131
  ## License
132
 
133
+ This model is released under the same license as the original HDLM codebase. Please refer to the [GitHub repository](https://github.com/ServiceNow/hdlm) for license details.