File size: 9,621 Bytes
3457594
549e18c
 
 
 
3457594
549e18c
 
3457594
549e18c
3457594
549e18c
3457594
549e18c
 
 
 
 
3457594
 
3a8f168
 
3457594
549e18c
3457594
6e8fa37
 
 
 
 
549e18c
 
 
3457594
549e18c
 
 
 
3457594
6e8fa37
 
 
 
 
 
 
 
 
 
 
 
 
549e18c
3457594
549e18c
 
 
 
 
 
 
3457594
549e18c
 
3457594
549e18c
 
 
 
3457594
549e18c
 
 
 
 
3457594
549e18c
 
 
3457594
6e8fa37
 
 
 
 
 
 
 
 
 
 
 
 
991eeb8
 
6e8fa37
 
 
549e18c
3457594
549e18c
3457594
549e18c
 
 
3457594
549e18c
 
 
 
3457594
549e18c
 
 
3457594
549e18c
3457594
549e18c
 
 
 
 
 
 
 
 
 
 
3457594
549e18c
3457594
549e18c
 
3457594
 
549e18c
 
 
 
 
3457594
 
549e18c
 
 
 
3457594
549e18c
d97f3d3
549e18c
 
 
 
 
d97f3d3
6e8fa37
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
---
license: other
library_name: custom
inference: false
pipeline_tag: image-text-to-text
tags:
- multimodal
- vision-language
- deepseek-v4
- moonvit
- nvfp4
- fp8
- blackwell
- sglang
base_model:
- MJPansa/DeepSeek-V4-Flash-0731-NVFP4
- deepseek-ai/DeepSeek-V4-Flash-0731
- moonshotai/Kimi-K2.6
---

# DeepSeek V4 Flash 0731 Vision (NVFP4)


![DeepSeek V4 Flash with MoonViT](deepseek-vision-improved.gif)

**DeepSeek V4 Flash 0731 with sight.** This development checkpoint connects
DeepSeek's reasoning and agentic backbone to the MoonViT vision encoder from
[Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6) through WebBrain's
trained, routing-aware PatchMerger projector.

This is a complete copy of MJPansa's source-preserving 0731 NVFP4 conversion
with WebBrain's frozen MoonViT tower, trained PatchMerger projector, routing
bridge, and custom SGLang integration added in the same repository.

The text checkpoint is mixed precision: routed expert weights are NVFP4 while
attention, shared experts, head, MTP, and other ignored paths preserve the
source 0731 formats described by `config.json`, `hf_quant_config.json`, and
`conversion-receipt.json`. It is not an all-NVFP4 checkpoint.

## Why vision at WebBrain

At [WebBrain](https://www.webbrain.one), we build browser agents that need to
understand the visual state of the web—not just extracted text. Screenshots,
charts, dashboards, rich editors, and the location and appearance of controls
are part of real browser work, so vision is a practical product requirement.

In our
[American–Chinese open-model frontier benchmark](https://www.webbrain.one/blog/american-chinese-open-model-frontier-gap-benchmark),
DeepSeek V4 Flash stood out as a strong and cost-efficient model, while the
upstream checkpoint remains text-only. This project attaches a MoonViT vision
bridge while keeping the language backbone and vision tower frozen.

## Contents

| Component | Provenance | Status |
|---|---|---|
| Text checkpoint | [`MJPansa/DeepSeek-V4-Flash-0731-NVFP4`](https://huggingface.co/MJPansa/DeepSeek-V4-Flash-0731-NVFP4) at [`64d64cd89bc63a66aa46506da89d7821f7491c62`](https://huggingface.co/MJPansa/DeepSeek-V4-Flash-0731-NVFP4/commit/64d64cd89bc63a66aa46506da89d7821f7491c62) | 48 shards; copied server-side without tensor changes |
| Conversion source | [`deepseek-ai/DeepSeek-V4-Flash-0731`](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) at `9e165c30e2704aec5d9d593cce3eebd58bbef1cb` | 0731 DSpark/MTP tensors preserved |
| Vision tower | MoonViT-3d extracted from [`moonshotai/Kimi-K2.6`](https://huggingface.co/moonshotai/Kimi-K2.6) at `7eb5002f6aadc958aed6a9177b7ed26bb94011bb` | 329 BF16 tensors; frozen |
| Projector | WebBrain PatchMerger, trained with the text backbone and tower frozen | 6 BF16 tensors; 40,119,040 parameters |
| Serving | External DeepSeek/MoonViT SGLang model and processor | Experimental; source patch required |

The added components are byte-identical to the published WebBrain vision
artifacts:

| File | Size | SHA-256 |
|---|---:|---|
| `vision_tower.safetensors` | 833,765,768 bytes | `1382c41f1a4afc91791ade630e2b1e1cef68cc5a1e09668a45970a5d5e1b8f15` |
| `mm_projector.safetensors` | 80,238,568 bytes | `7024d9d5c9714c7abbc09abda015f083b7d7b107745eb78879f019bf4721577a` |

The initial mirror commit, `3457594532d890f8a46c18dda72d6a85c94ab418`,
matched all 78 source paths, sizes, Git blobs, and LFS SHA-256 values
(175,573,280,882 bytes). The final package intentionally replaces the model
card and augments `config.json` and `model.safetensors.index.json` so the vision
components can be discovered. The 48 text shards remain unchanged.

See [`VISION_ADAPTER_MANIFEST.json`](VISION_ADAPTER_MANIFEST.json) and the
preserved [`conversion-receipt.json`](conversion-receipt.json) for
machine-readable provenance.

## Build status

- [x] Pin and exactly mirror the source-preserving 0731 NVFP4 conversion.
- [x] Preserve all 48 text shards, the conversion receipt, and 0731 DSpark/MTP
  tensors.
- [x] Preserve mixed-precision metadata: routed experts are NVFP4 while ignored
  attention, shared-expert, head, and MTP paths retain their source formats.
- [x] Add and fingerprint-verify the frozen BF16 MoonViT tower.
- [x] Add and fingerprint-verify the trained BF16 PatchMerger projector.
- [x] Add the routing palette, external processor/model package, source patch,
  launch wrapper, and regression tests.
- [x] Augment the source config and index without changing upstream text tensor
  mappings or quantization metadata.
- [x] Pass a full-model GPU loader/startup gate and live image-generation smoke
  test with the established deployment path.
- [ ] Run fresh text-parity and live image-generation smoke tests on the target
  Blackwell hardware and SGLang image.

## Integration details

The projector performs:

```text
LayerNorm -> 2x2 patch merge -> Linear(4608, 4608) -> GELU -> Linear(4608, 4096)
```

Text token routing IDs are preserved. Image positions receive deterministic
routing IDs from the checked-in 64-ID palette. The integration currently
supports one literal `<image>` marker and targets at most 512 merged image
tokens.

This is not a stock Transformers or stock SGLang multimodal architecture. Use
the checked-in external model/processor package and the SGLang revision pinned
in [`docs/SGLANG_DEPLOYMENT.md`](docs/SGLANG_DEPLOYMENT.md).

## Validation status

- Exact source mirror metadata was verified before the vision overlay.
- Both added safetensors files were verified by LFS SHA-256 after publication.
- The source config's architecture, 0731 DSpark fields, mixed FP8/NVFP4
  quantization metadata, and per-layer routed-expert map are preserved; only
  vision metadata was added.
- The index still references the same 48 text shards and additionally maps the
  335 vision/projector tensor names.
- **A full GPU loader/startup and image-generation smoke test has not yet been
  rerun for this assembled 0731 checkpoint.** Earlier validation of the same
  vision adapter and SGLang glue with another NVFP4 text package is not claimed
  as validation here.

## Usage

Download the complete repository, then point the included wrapper at that local
snapshot:

```bash
export DEEPSEEK_VISION_MODEL_PATH=/models/DeepSeek-V4-Flash-0731-Vision-NVFP4
export DEEPSEEK_VISION_PYTHONPATH="$DEEPSEEK_VISION_MODEL_PATH/sglang_ext"
export DEEPSEEK_VISION_KERNEL_PROFILE=blackwell-native
export DEEPSEEK_VISION_TP=4
scripts/launch_sglang_moonvit.sh
```

Treat the Blackwell profile as a starting configuration, not a validation
claim. Begin with SGLang's native `/generate` endpoint as documented in the
deployment guide. OpenAI chat image parts, production concurrency, CUDA graphs,
and broad quality behavior remain unvalidated.

## Licenses

The converted text files retain the terms and notices from their source. The
MoonViT-derived tower is covered by `LICENSE_KIMI_K2.6`. The WebBrain projector
and integration source are provided under the terms included with this
repository. Downstream users are responsible for satisfying all applicable
licenses.

## Method credit

The overall construction and model-card approach was inspired by
[Baseten's GLM-5.2-Vision-NVFP4](https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4):
keep the text backbone and MoonViT tower frozen, train a compact PatchMerger
projector between them, and publish provenance and hardware constraints
explicitly. Credit to the Baseten team for demonstrating this practical recipe.

This project adds a DeepSeek-specific routing bridge so mixed text/image
embeddings preserve hash-routed text behavior. It does not reuse Baseten model
weights, benchmark results, or deployment artifacts.

## Acknowledgements

Built on
[DeepSeek AI's DeepSeek V4 Flash 0731](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731),
[MJPansa's source-preserving 0731 NVFP4 conversion](https://huggingface.co/MJPansa/DeepSeek-V4-Flash-0731-NVFP4),
and [Moonshot AI's Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6),
with the vision-attachment method inspired by
[Baseten's GLM-5.2-Vision-NVFP4](https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4).
These teams were not involved in this experimental development checkpoint;
please do not direct issues with this repository to them.

## Want this model on your inference provider?

Ask your inference provider—such as
[OpenRouter](https://openrouter.ai/) or another managed inference service—to
deploy this exact repository with its multimodal processor and serving plugin.
Deploying only the upstream text model will not enable image input. Providers
should complete the fresh 0731 GPU validation gates above before offering it as
a production endpoint.

## Experimental status, roadmap, and get involved

> [!CAUTION]
> **Experimental vision checkpoint.** Repository assembly, provenance, hashes,
> and source preservation have been verified, but a fresh full GPU loader and
> image-generation smoke test has not been run for this 0731 package. This is
> not a production-readiness or quality claim. Fine-grained OCR, small-object
> and control identification, GUI grounding, hallucination calibration,
> concurrency, and broader hardware coverage still need evaluation. Do not use
> this model as the sole decision source for safety-critical automation. The
> roadmap includes larger and more diverse datasets, higher-resolution OCR/UI
> examples, and broader parameter-efficient tuning. Interested in contributing
> evaluation or training data, sponsoring compute, or working with us as a
> design partner? [Tell us here](https://forms.gle/bNoeJ6cvLYQ4VgKd7).