Ya-Lin-Zhang commited on
Commit
86d20ac
·
verified ·
1 Parent(s): 9f6c565

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +7 -10
README.md CHANGED
@@ -16,7 +16,7 @@ Key highlights of the model are summarized below:
16
  + **Remarkable Efficiency & Performance:** Engineered for speed, compute efficiency, and production deployment, Ling-3.0-flash delivers class-defying performance against both larger SOTA competitors and previous-generation flagships. Activating only 5.1B parameters per token, it provides impressive reasoning, instruction following, and long-context capabilities to empower complex agentic workflows in production environments.
17
  + **Comprehensive Agentic Evolution:** Tailored for real-world productivity workflows, the model incorporates 10,000+ interactive training environments to achieve end-to-end closed-loop execution across Coding, General, and Deep Research Agent tasks. It natively integrates the SGLang HiCache + Mooncake hierarchical caching architecture (featuring physical dual-pools and a cluster-shared L3 cache), eliminating redundant recomputation during long-horizon interactions and reducing Time to First Token (TTFT) by 60% to over 80% in long-input scenarios.
18
 
19
- <!-- 这是一张图片,ocr 内容为:SWE-BENCH MULTILINGUAL TERMINAL-BENCH 2.1 SWE-BENCHPRO TAU3-BANKING-AA LL088000LL 65 6 71.2 76.5 75.9 AP 28.0 60 0 72.4 73.3 56.6 56.2 56.3 71.0 53.9 70 22.9 55 57.0 50 47.9 48.3 60 56.7 14.6 45 11.3 71.3 39.3 39.0 70.1 50 8.9 40 42.7 40 34.1 35 30 30 WIDESEARCH MCP-ATLAS SKILLSBENCH BROWSECOMP 品质88品质品等学导新品 品民品品导品只只口 豆豆复品品导导导 品品&品品定品名品 6 4 79.5 米 0 75.2 73.6 74.4 70.2 69.0 82.0 66.7 6 65.5 62.2 53.5 75.8 61.2 73.2 71.7 74.0 55.2 53.6 52.6 49.4 G 20.3 19.5 71.9 31.3 MULTI-AGENT IFBENCH SYSBENCH MRCR-256K MULTI-IF 100 雪饼 110 110 87.7 89.3 79.2 06 82.9 84.6 86.2 84.6 84.8 75.7 82.3 93.6 91.4 93.9 90793.3949 72.6 84.3 90 品8888元 81.1 86.5 86.2 69.0 70 67.3 70 60 56.6 05 30 0 30 10 ULIO RING-2.6-1T(XHIGH) STEP-3.7-FLASH(HIGH) LING-3.0-FLASH MINIMAX-M2.7 WALBAN DEEPSEEK-V4-FLASH-PREVIEW(MAX) CLAUDE-SONNET-4.6(MAX) GPT-5.4-MINI(HIGH) NEMOTRON-3-SUPER-120B-A12B NOTE:THINKING MODE IS ENABLED BY DEFAULT. -->
20
  ![](https://intranetproxy.alipay.com/skylark/lark/0/2026/png/23157180/1785831264180-d6ca4404-acef-4424-84db-fbc5a4c6db5f.png)
21
 
22
  ## Model Overview
@@ -24,7 +24,7 @@ The model summary information and architecture diagram are as follows:
24
 
25
  | Architecture | Hybrid-linear MoE |
26
  | --- | --- |
27
- | Parameter Scale | Totoal 124B, Activated 5.1B |
28
  | Transformer Layers | 35 KDA + 7 Gated MLA (5:1) |
29
  | Number of Dense Layers | 2 |
30
  | Number of Routed Experts | 512 |
@@ -37,28 +37,25 @@ The model summary information and architecture diagram are as follows:
37
  | Vocabulary Size | 157184 |
38
  | Context Training Schedule | 8K -> 32K -> 256K |
39
 
40
-
41
-
42
-
43
- <!-- 这是一张图片,ocr 内容为:LING-3.0-FLASH ARCHITECTURE TRAININGOBJECTIVE:NEXT-TOKENPRE TOKEN PREDICTION AND MULTI-TOKEN PREDICTION (MTP) VOCABULARY O:SIGMOID GATE OUTPUT LINEAR OUTPUT LAYER SIZE OF 157K PROJECTION FINAL RMSNORM MULTI-HEAD LATENT ATTENTION K E512A8+1SHARED MOE ROPE ROPE ERT,ALF-LB EXPERT RMSNORM LINEAR LINEAR RMSNORM RMSNORM SUPPORTED LINEAR LINEAR LINEAR GATED MLA ROPE CONTENTLENGTH-- OF 1M TOKENS RMSNORM O:SIGMOID GATE OUTPUT P:SOFTPLUS GATE PROJECTION 8:SWISH FUNC RMSNORM MOE KIMI DELTA ATTENTION RMSNORM K V L2 NORM LINEAR TIME 8 KDA COMPLEXITY RMSNORM CONV CONV MLP LINEAR TOKEN EMBEDDING LAYER 7 GROUPS LINEAR LINEAR FIRST 2 BLOCKS EMBEDDING TOKENIZED TEXT USE DENSE FFN DIMENSION OF 2,560 个 INSTEAD OF MOE SAMPLE INPUT TEXT -->
44
  ![](https://intranetproxy.alipay.com/skylark/lark/0/2026/png/23157180/1785822388609-79c06c50-2ea9-40d9-888b-0ff4072a724c.png)
45
 
46
  ## Evaluation
47
  We have conducted a comprehensive evaluation of Ling-3.0-flash across multiple authoritative benchmarks. **Ling-3.0-flash** performs strongly on representative code/agent benchmarks such as **SWE-Bench Pro, SWE-Bench Multilingual, Tau3-banking-AA**, **MCP-Atlas** and **SkillsBench, etc**. In practice, Ling-3.0-flash delivers a strong user experience across frameworks including **Claude Code**,**Kilo Code**,**Qwen Code**,**Hermes Agent**,and **OpenClaw**, etc. Beyond agentic tasks, Ling-3.0-flash also delivers strong performance across **general knowledge**,**mathematical reasoning**,**instruction following**,and **long-context understanding**.
48
 
49
- <!-- 这是一张图片,ocr 内容为:DEEPSEEK CLAUDE- NEMOTRON- LING-3.0 GPT-5.4 STEP-3.7- V4-FLASH- RING-2.6-1T MINIMAX- 3-SUPER FLASH (HIGH) MINI (HIGH) FLASH (XHIGH) 120B-A12B (MAX) (MAX) 120B-A12B 124B-A5.1B 284B-A13B 198B-A11B 230B-A10B 1T-A63B SIZE CODING AGENT SWE-BENCH PRO 47.9 34.1 48.3 56.2 56.3 52.6 53.9 56.6 SWE-BENCH MULTILINGUAL 72.4 71.0 42.7 56.7 724 76.5 73.3 75.9 57.0 TERMINAL-BENCH 2.1 39.3 55.8 55.0 71.2 62.0 43.1 39.0 59.2 77.0 65.1 55.8 ARTIFACTSBENCH 66.8 64.0 51.6 68.7 25.3 19.2 MINIAPPENCH 28.0 14.8 46.3 20.7 58.8 5.8 52.2 ANTSWEBENCH 48.6 46.4 GENERAL AGENT 8.9 11.3 28.0 TAU3-BANKING-AA 11.3 10.1 22.9 14.6 30.5 55.2 53.6 69.0 65.5 MCP-ATLAS 52.6 61.2 49.4 66.7 53.5 24.9 20.3 44.8 54.4 28.4 44.8 11.9 SKILLSBENCH BFCL-V4 59.5 73.0 68.3 73.1 65.4 63.6 60.6 64.8 GDPVAL V2-AA 920 1377 1189 1017 1107 699 1159 SEARCH AGENT WIDESEARCH 75.2 70.2 62.2 73.6 19.5 74.4 56.8 79.5 72.2(W/ CTX) 74.0(W/CTX) BROWSECOMP 73.2 71.7 31.3 75.8 76.3 82.0(MA) 82.1(MA) DRACO 71.3 75.8 61.3 66.8 70.4 INSTRUCTION FOLLOWING 75.7 74.5 IFBENCH 72.6 67.3 79.2 69.0 44.6 56.6 91.4 93.6 93.3 SYSBENCH 86.5 93.9 90.7 86.2 94.9 77.3 71.3 LIFEBENCH 69.2 72.5 71.8 74.1 66.9 60.2 REASONING 91.7 94.2 AIME26 93.2 95.0 95.8 92.9 96.5 94.4 HMMT-FEB26 87.0 71.9 93.5 85.6 94.8 87.9 84.9 83.9 83.7 74.5 87.0 IMO-ANSWERBENCH 77.0 82.1 86.1 66.9 22.7 HLE 19.9 20.6 18.3 18.3 34.8 28.1 30.0 LIVECODEBENCH 82.8 78.1 80.8 87.0 75.6 83.7 78.7 91.6 (2408-2505) LONG CONTEXT & MULTI-TURN DIALOGUE 39.2 MRCR 128K 27.7 90.1 56.1 40.8 90.8 92.5 88.5 92.7 81.1 MRCR 256K 25.7 76.5 50.5 35.8 84.3 70.7 AA-LCR 65.1 64.3 68.7 63.7 63.0 58.3 63.4 89.3 MULTI-IF 86.2 87.7 82.9 84.6 84.8 82.3 84.6 -->
50
  ![](https://intranetproxy.alipay.com/skylark/lark/0/2026/png/23157180/1785843565924-387dbbd7-90f8-4241-a58d-664cc5e6b486.png)
51
 
52
  > + Thinking mode is enabled by default. Unless otherwise specified, the default parameters for Ling-3.0-flash are as follows: `temperature=0.6, top_p=0.95, top_k=20`.
53
- > + SWE-Bench Series:Evaluated using OpenHands as the agent harness with tailored prompts. Decoding uses `temperature=0.6, top_p=0.95, max_new_tokens=32K`, with a 256K context window.
54
  > + Terminal-Bench 2.1: Evaluated under the Artificial Analysis (AA) protocol using the default Terminus 2 harness, a unified 2-hour timeout, the provided JSON parser in preserve-thinking mode, and 3 runs per task (mean). Decoding uses `temperature=0.6, top_p=1.0, max_new_tokens=32K`, with a 256K context window.
55
  > + MiniAppBench: A 500-task coding benchmark evaluating whether models can turn a single user request into complete, usable interactive HTML apps in real-world application-generation scenarios. Evaluated with `temperature=1.0, top_p=1.0, max_tokens=128K`.
56
  > + AntSWEBench: AntSWEBench is an internally used software engineering benchmark that covers mainstream programming languages such as Java, JavaScript, and Python, including various development scenarios like new feature, bug fix, and code refactoring.
57
  > + Tau3-banking-AA: Aligned with the AA leaderboard, utilizing GPT-5.4-mini (medium reasoning) for both the user simulator and the natural-language assertion judge.
58
  > + MCP-Atlas: Evaluated on the 500-task public set using the official v1 harness with a 20-turn limit and Gemini-2.5-Pro as the claim-coverage judger.
59
  > + SkillsBench: Evaluated via kilo-code on 87 tasks (excluding external API-dependent tasks), averaged over 3 runs.
60
- > + GDPval v2-AA : Evaluated on the public 220-task benchmark using the official Stirrup harness, with a 250-turn limit and a 5-hour timeout.
61
- > + SearchagentFor all search‑agent tasks, evaluations are performed using an internal harness. The basic ReAct paradigm is adopted for single-agent evaluation, while a multi-agent setup is employed for BrowseComp. The reported metric is the average pass@1.
62
  > - WideSearch: Evaluated using the official prompt and the official judge model GPT-4.1 on the corrected version of the dataset.
63
  > - Draco: Scored based on official rubrics per question, with the final score calculated as the average across all questions using Claude Opus 4.6 as the scoring model.
64
  > - BrowseComp (Single-Agent): Evaluated using a resume strategy for context management: once the context reaches a 64K-token threshold, the trajectory is summarized, the original history is discarded, and execution is resumed from the summary.
 
16
  + **Remarkable Efficiency & Performance:** Engineered for speed, compute efficiency, and production deployment, Ling-3.0-flash delivers class-defying performance against both larger SOTA competitors and previous-generation flagships. Activating only 5.1B parameters per token, it provides impressive reasoning, instruction following, and long-context capabilities to empower complex agentic workflows in production environments.
17
  + **Comprehensive Agentic Evolution:** Tailored for real-world productivity workflows, the model incorporates 10,000+ interactive training environments to achieve end-to-end closed-loop execution across Coding, General, and Deep Research Agent tasks. It natively integrates the SGLang HiCache + Mooncake hierarchical caching architecture (featuring physical dual-pools and a cluster-shared L3 cache), eliminating redundant recomputation during long-horizon interactions and reducing Time to First Token (TTFT) by 60% to over 80% in long-input scenarios.
18
 
19
+ <!-- Benchmark comparison chart across models -->
20
  ![](https://intranetproxy.alipay.com/skylark/lark/0/2026/png/23157180/1785831264180-d6ca4404-acef-4424-84db-fbc5a4c6db5f.png)
21
 
22
  ## Model Overview
 
24
 
25
  | Architecture | Hybrid-linear MoE |
26
  | --- | --- |
27
+ | Parameter Scale | Total 124B, Activated 5.1B |
28
  | Transformer Layers | 35 KDA + 7 Gated MLA (5:1) |
29
  | Number of Dense Layers | 2 |
30
  | Number of Routed Experts | 512 |
 
37
  | Vocabulary Size | 157184 |
38
  | Context Training Schedule | 8K -> 32K -> 256K |
39
 
40
+ <!-- Ling-3.0-flash architecture diagram -->
 
 
 
41
  ![](https://intranetproxy.alipay.com/skylark/lark/0/2026/png/23157180/1785822388609-79c06c50-2ea9-40d9-888b-0ff4072a724c.png)
42
 
43
  ## Evaluation
44
  We have conducted a comprehensive evaluation of Ling-3.0-flash across multiple authoritative benchmarks. **Ling-3.0-flash** performs strongly on representative code/agent benchmarks such as **SWE-Bench Pro, SWE-Bench Multilingual, Tau3-banking-AA**, **MCP-Atlas** and **SkillsBench, etc**. In practice, Ling-3.0-flash delivers a strong user experience across frameworks including **Claude Code**,**Kilo Code**,**Qwen Code**,**Hermes Agent**,and **OpenClaw**, etc. Beyond agentic tasks, Ling-3.0-flash also delivers strong performance across **general knowledge**,**mathematical reasoning**,**instruction following**,and **long-context understanding**.
45
 
46
+ <!-- Benchmark evaluation results comparison chart -->
47
  ![](https://intranetproxy.alipay.com/skylark/lark/0/2026/png/23157180/1785843565924-387dbbd7-90f8-4241-a58d-664cc5e6b486.png)
48
 
49
  > + Thinking mode is enabled by default. Unless otherwise specified, the default parameters for Ling-3.0-flash are as follows: `temperature=0.6, top_p=0.95, top_k=20`.
50
+ > + SWE-Bench Series: Evaluated using OpenHands as the agent harness with tailored prompts. Decoding uses `temperature=0.6, top_p=0.95, max_new_tokens=32K`, with a 256K context window.
51
  > + Terminal-Bench 2.1: Evaluated under the Artificial Analysis (AA) protocol using the default Terminus 2 harness, a unified 2-hour timeout, the provided JSON parser in preserve-thinking mode, and 3 runs per task (mean). Decoding uses `temperature=0.6, top_p=1.0, max_new_tokens=32K`, with a 256K context window.
52
  > + MiniAppBench: A 500-task coding benchmark evaluating whether models can turn a single user request into complete, usable interactive HTML apps in real-world application-generation scenarios. Evaluated with `temperature=1.0, top_p=1.0, max_tokens=128K`.
53
  > + AntSWEBench: AntSWEBench is an internally used software engineering benchmark that covers mainstream programming languages such as Java, JavaScript, and Python, including various development scenarios like new feature, bug fix, and code refactoring.
54
  > + Tau3-banking-AA: Aligned with the AA leaderboard, utilizing GPT-5.4-mini (medium reasoning) for both the user simulator and the natural-language assertion judge.
55
  > + MCP-Atlas: Evaluated on the 500-task public set using the official v1 harness with a 20-turn limit and Gemini-2.5-Pro as the claim-coverage judger.
56
  > + SkillsBench: Evaluated via kilo-code on 87 tasks (excluding external API-dependent tasks), averaged over 3 runs.
57
+ > + GDPval v2-AA: Evaluated on the public 220-task benchmark using the official Stirrup harness, with a 250-turn limit and a 5-hour timeout.
58
+ > + Search-agent: For all search‑agent tasks, evaluations are performed using an internal harness. The basic ReAct paradigm is adopted for single-agent evaluation, while a multi-agent setup is employed for BrowseComp. The reported metric is the average pass@1.
59
  > - WideSearch: Evaluated using the official prompt and the official judge model GPT-4.1 on the corrected version of the dataset.
60
  > - Draco: Scored based on official rubrics per question, with the final score calculated as the average across all questions using Claude Opus 4.6 as the scoring model.
61
  > - BrowseComp (Single-Agent): Evaluated using a resume strategy for context management: once the context reaches a 64K-token threshold, the trajectory is summarized, the original history is discarded, and execution is resumed from the summary.