Edge Rewrite
Jump to content

DeepSeek

From Wikipedia, the free encyclopedia

Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd.
Native name
杭州深度求索人工智能基础技术研究有限公司
TypePrivate
IndustryInformation technology
Artificial intelligence
Founded17 July 2023; 3 years ago (2023-07-17)[1]
Founder
HeadquartersHangzhou, Zhejiang,
China
Key people
  • Liang Wenfeng (CEO)
ProductsDeepSeek
OwnerHigh-Flyer
Number of employees
160 (2025)[2]
Websitedeepseek.com

Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd.,[3][4][5][a] doing business as DeepSeek,[b] is a Chinese artificial intelligence (AI) company that develops open weights large language models (LLMs). Based in Hangzhou, Zhejiang, DeepSeek is owned and funded by High-Flyer, a Chinese hedge fund.

Background

[edit]

DeepSeek was founded in July 2023 by Liang Wenfeng.[7][8][9] The company launched an eponymous chatbot alongside its DeepSeek-R1 model in January 2025.

DeepSeek's models are open-weight, meaning that the exact parameters are openly shared, but the training data is not openly licensed.[10][11] Since January 2025, DeepSeek has released its models under free and open-source software licenses.[12] The company reportedly recruits AI researchers from top Chinese universities and hires from outside traditional computer science fields to broaden its models' knowledge and capabilities.[13][14] DeepSeek is considered more closely linked to China's defense industry due to previous affiliations of its researchers and its chatbot's adoption for non-combat roles by the People's Liberation Army since March 2025.[15]

History

[edit]

Pre-founding years

[edit]

In June 2015, the hedge fund High-Flyer was co-founded by Liang Wenfeng.[16] By 2021, Liang had started buying large quantities of Nvidia GPUs for an AI project,[17] reportedly obtaining 10,000 Nvidia A100 GPUs before the United States restricted chip sales to China. Fewer than 5 Chinese tech companies owned more than 10,000 GPUs.[18][19]

Founding

[edit]

On 14 April 2023,[20] High-Flyer announced the launch of an artificial general intelligence (AGI) research lab, stating that the new lab would focus on developing AI tools unrelated to the firm's financial business.[21][22] Two months later, on 17 July 2023,[1] that lab was spun off into an independent company. That lab was named DeepSeek, with High-Flyer as its principal investor and backer.[18][22][23] Initially, venture capital investors were reluctant to provide funding, as they considered it unlikely that the venture would be able to quickly generate an exit (an ownership stake in an investment).[18]

Company operation

[edit]

DeepSeek is headquartered in Hangzhou, Zhejiang, and is owned and funded by High-Flyer. Its co-founder, Liang Wenfeng, serves as CEO. As of May 2024, Liang personally held an 84% stake in DeepSeek through two shell corporations.[note 1][24] In February 2026, Anthropic accused DeepSeek of using thousands of fraudulent accounts to generate millions of conversations with Claude to train its own LLMs.[25] In April 2026, investors began speaking with DeepSeek for a $300 million funding round, which would bring DeepSeek to a total valuation of $10 billion.[26][needs update] DeepSeek completed a May 2026 Series A round receiving US$7 billion, reaching a post-money valuation of US$52 billion.[27] In July 2026, Bloomberg and the Financial Times reported the company had begun preparations for an IPO that would see it list as soon as 2027.[28] The same month, it began another discussion targeting a US$70 billion pre-money valuation.[29]

Due to its open weight framework, DeepSeek models have been hosted natively by cloud providers including Microsoft Azure and Perplexity AI.[30]

The DeepSeek login page following a cyberattack around its 21 January 2025 launch

Strategy

[edit]

DeepSeek has stated that it focuses on research and does not have immediate plans for commercialization.[31] This posture also means it can skirt certain provisions of China's AI regulations aimed at consumer-facing technologies.[13][needs update][original research?]

DeepSeek operates with a relaxed, academic research-style workplace culture that deliberately avoids China’s "996" overtime practices. Founder Liang Wenfeng structured the company to eliminate rigid Key Performance Indicators (KPIs) and mandatory overtime, allowing significant time for unstructured research and creative exploration.[32]

DeepSeek's hiring approach emphasizes skills over lengthy work experience, resulting in many hires fresh out of university.[13][22] The company likewise recruits individuals without computer science backgrounds to expand the range of expertise incorporated into the models, for instance in poetry or advanced mathematics.[13][14] According to The New York Times, dozens of DeepSeek researchers have or have previously had affiliations with People's Liberation Army laboratories and the Seven Sons of National Defence.[33] Since 2025, its chatbot was adopted in non-combat roles by the People's Liberation Army, including hospitals, the People's Armed Police paramilitary, and national mobilization organizations.[15]

Due to U.S. chip restrictions, DeepSeek has continuously refined its algorithms to maximize computational efficiency, leveraging older hardware and reducing energy consumption.[34]:19

DeepSeek also expanded into Africa, offering more affordable and less power-hungry AI solutions. The company has bolstered African language models and generated a number of startups, for example in Nairobi. Along with Huawei's storage and cloud computing services, the impact on the tech scene in sub-Saharan Africa is considerable. DeepSeek offers local data sovereignty and more flexibility compared to Western AI platforms.[35]

Training framework

[edit]

High-Flyer/DeepSeek had operated at least two primary computing clusters: Fire-Flyer and Fire-Flyer 2. Fire-Flyer 1 was constructed in 2019 and was retired after 1.5 years of operation.[citation needed] Fire-Flyer 2 is still in operation as of 2025.[citation needed] Fire-Flyer 2 consists of co-designed software and hardware architecture.[citation needed] On the hardware side, Nvidia GPUs use 200 Gbps interconnects.[citation needed] The cluster is divided into two zones, and the platform supports cross-zone tasks.[citation needed] The network topology consisted of two fat trees, chosen for their high bisection bandwidth. On the software side, there are:[36]

  • 3FS (Fire-Flyer File System): A distributed parallel file system designed for asynchronous random reads. It uses Direct I/O and RDMA Read. In contrast to standard Buffered I/O, Direct I/O does not cache data. Caching is useless in this case, since each piece of data read is random and is not reused.[37][38]
  • hfreduce: Library for asynchronous communication, originally designed to replace Nvidia Collective Communication Library (NCCL).[39] It is mainly used for allreduce, especially of gradients during backpropagation.[citation needed] It runs asynchronously on the CPU to avoid blocking kernels on the GPU.[36] It uses two-tree broadcast like NCCL.[39]
  • hfai.nn: Software library of commonly used operators for neural network training, similar to torch.nn in PyTorch.[citation needed]
  • HaiScale Distributed Data Parallel (DDP): Parallel training library that implements various forms of parallelism such as data parallelism (DP), pipeline parallelism (PP), tensor parallelism (TP), experts parallelism (EP), fully sharded data parallel (FSDP) and zero redundancy optimizer (ZeRO). It is similar to PyTorch DDP, which uses NCCL on the backend.
  • HAI Platform: Various applications such as task scheduling, fault handling, and disaster recovery.[40]

As of 2022, Fire-Flyer 2 had 5,000 PCIe A100 GPUs in 625 nodes, each containing 8 GPUs.[39]

DeepSeek Harness

[edit]

DeepSeek Harness is an open-source AI agent harness developed by DeepSeek. It provides an environment for building and running AI agents, with components including models, tools, skills, sessions, sandboxes, storage, scheduling, and user interfaces. The software uses Cordis plugin architecture and is distributed under the MIT License. As of August 2026, DeepSeek Harness is available as a developer preview.[41]

Plugin architecture

[edit]

DeepSeek Harness uses an "everything-is-a-plugin" design philosophy powered by the Cordis framework.[41] This allows major agent components to be added, removed, or reloaded while the software is running.[41][42] The modular design enables developers to switch between third-party providers or alternative model configurations without changing the rest of the system.[42]

Model releases

[edit]

On 21 August 2025, DeepSeek released DeepSeek V3.1 under the MIT License.[43] This model features a hybrid architecture with thinking and non-thinking modes. It also surpasses prior models like V3 and R1, by over 40% on certain benchmarks like SWE-bench and Terminal-bench.[44] It was updated to V3.1-Terminus on 22 September 2025, which fixed multiple language artifacts and improved in intelligence in general.[45]

DeepSeek V3.2-Exp was released on 29 September 2025. It uses DeepSeek Sparse Attention, a more efficient attention mechanism based on previous research published in February. It was based on DeepSeek V3.1 Terminus.[46][47] DeepSeek-V3.2 was released on 1 December 2025, alongside a DeepSeek-V3.2-Speciale variant that focused on reasoning.[48][49]

R1 series

[edit]

On 20 January 2025, DeepSeek launched the DeepSeek-R1 model, which was free for iOS and Android. By 27 January, DeepSeek surpassed ChatGPT as the most downloaded freeware app on the iOS App Store in the United States,[14] triggering an 18% drop in Nvidia's share price.[50][51]

On 28 May 2025, DeepSeek released DeepSeek-R1-0528 under the MIT License, which was a minor update upon R1. Its scores have also exceeded both DeepSeek R1 (January) and DeepSeek V3-0324 in certain benchmarks.[52]

V4 series

[edit]

On 24 April 2026, DeepSeek released a preview of its V4 series, including the 284-billion parameter DeepSeek-V4-Flash and the 1.6-trillion parameter DeepSeek-V4-Pro. Both feature a one million token context window, under the MIT License.[53][54][55] V4 has been adopted by semiconductor manufacturers such as Huawei and Cambricon Technologies.[56] The official versions of DeepSeek V4-Flash and V4-Pro were released on 31 July and 13 August, respectively.[57][58] V4.1-Flash, which was made available on 10 September 2026, introduced a "causal encoder-decoder" system and had lower memory usage compared to previous versions, according to DeepSeek.[59][60]

List of models

[edit]
Major versions of DeepSeek models. SFT stands for supervised finetuning, which allows AI labs to precisely adjust an AI model's behavior.
Major versions Release date Status Major variants License Remarks
DeepSeek-Coder November 2, 2023 Discontinued Base (pretrained)
Instruct (with instruction-finetuned)
Open-weights (DeepSeek) The architecture is essentially the same as Llama.[61][original research?]
DeepSeek-LLM November 29, 2023 Discontinued Base
Chat (with SFT)
DeepSeek-MoE January 9, 2024 Discontinued Base
Chat
Developed a variant of MoE.[62]
DeepSeek-Math April 2024 Discontinued Base Initialized with DS-Coder-Base-v1.5[63]
Instruct (with SFT) [64]
RL (using a process reward model) Developed Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO).[65]
DeepSeek-V2 May 2024 Discontinued DeepSeek-V2, DeepSeek-V2-Chat
DeepSeek-V2-Lite, DeepSeek-V2-Lite-Chat
DeepSeek-Coder-V2
DeepSeek-V2.5
Developed multi-head latent attention (MLA). Also used MoE. Implemented KV caching.[66]
DeepSeek-V3 December 2024 Active DeepSeek-V3-Base
DeepSeek-V3 (a chat model)
The architecture is essentially the same as V2. Updated on 24 March 2025.[67][original research?]
DeepSeek-Prover-V2 May 1, 2025 Active DeepSeek-Prover-V2-671B
DeepSeek-Prover-V2-7B
[68]
DeepSeek-VL2 December 13, 2024 Active [69]
DeepSeek-R1 November 20, 2024 Active DeepSeek-R1-Lite-Preview Proprietary Preview version, only accessed through API and a chat interface.
January 20, 2025 Active DeepSeek-R1
DeepSeek-R1-Zero
DeepSeek-R1-0528
MIT Initialized from DeepSeek-V3-Base and sharing the V3 architecture.[52]
Distilled models Initialized from other models, such as Llama, Qwen, etc. Distilled from data synthesized by R1 and R1-Zero.[70]
May 28, 2025 Active DeepSeek-R1-0528 N/A
DeepSeek-V3.1 August 21, 2025 Active DeepSeek-V3.1-Base
DeepSeek-V3.1 (a chat model)
Hybrid architecture (thinking and non-thinking modes available). Trained on over 800B additional tokens on top of V3.[71]
September 22, 2025 Active DeepSeek-V3.1-Terminus Reducing instances of mixed Chinese-English text and occasional abnormal characters on top of V3.1.[72]
DeepSeek-Math-V2 November 27, 2025 Active Apache 2.0 [73]
DeepSeek-V3.2 December 1, 2025 Active DeepSeek-V3.2
DeepSeek-V3.2-Speciale
MIT [48][49][74]
DeepSeek-V4 April 24, 2026 Active V4-Pro, V4-Flash Preview release on 24 April.[53][54][55] Official release of V4-Flash on 31 July,[57] and V4-Pro on 13 August[58]
DeepSeek-V4.1 September 10, 2026 Active V4.1-Flash [59][60][75]

Technical specifications of models

[edit]

DeepSeek's models are open-weight, which provides less freedom for modification than true open source software.[10][11][original research?]

[edit]

In the United States, the National Defense Authorization Act for Fiscal Year 2026 instructs the United States Secretary of Defense and the Director of National Intelligence to remove and exclude any artificial intelligence developed by DeepSeek or owned by HighFlyer from devices operated by the United States Department of Defense and United States Intelligence Community or their contractors. Artificial intelligence by DeepSeek is only allowed with explicit permissions for research, training, and evaluation or military activities supporting national security functions such as counterterrorism or counterintelligence.[76][77]

In Australia, the Department of Home Affairs issued a government-wide directive "to prevent the use or installation of DeepSeek products, applications and web services and where found remove all existing instances of DeepSeek products, applications and web services from all Australian Government systems and devices."[78]

See also

[edit]

Notes

[edit]
  1. Chinese: 杭州深度求索人工智能基础技术研究有限公司.[6] Sometimes simply referred to in English as Hangzhou DeepSeek Artificial Intelligence.
  2. Chinese: 深度求索; pinyin: Shēndù Qiúsuǒ
  1. 宁波程信柔兆企业管理咨询合伙企业(有限合伙) and 宁波程恩企业管理咨询合伙企业(有限合伙)

References

[edit]
  1. 1 2 DeepSeek突传消息. Sina Corporation. 1 February 2025. Retrieved 1 February 2025.
  2. Wu, Zijing (14 March 2025). "DeepSeek focuses on research over revenue in contrast to Silicon Valley". Financial Times. Retrieved 14 March 2025.
  3. "Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd". Bloomberg L.P.
  4. "DeepSeek Coder Model Service Agreement" (PDF), DeepSeek, 19 October 2023, archived (PDF) from the original on 21 February 2025, retrieved 11 February 2025
  5. "DeepSeek Coder Privacy Policy" (PDF). DeepSeek. Archived from the original (PDF) on 8 July 2025. Retrieved 19 February 2025.
  6. 全国互联网安全管理平台. beian.mps.gov.cn (in Chinese (China)). Ministry of Public Security of the People's Republic of China. Archived from the original on 9 February 2025. Retrieved 9 February 2025.
  7. Jiang, Ben (21 January 2025). "Beijing puts spotlight on China's new face of AI, DeepSeek's Liang Wenfeng". South China Morning Post. Archived from the original on 21 January 2025. Retrieved 4 March 2025.
  8. Baptista, Eduardo (28 January 2025). "Who is Liang Wenfeng, the founder of DeepSeek?". Reuters. Archived from the original on 19 February 2025. Retrieved 4 March 2025.
  9. "Behind DeepSeek lies a dazzling Chinese university". The Economist. ISSN 0013-0613. Retrieved 5 March 2025.{{cite news}}: CS1 maint: deprecated archival service (link)
  10. 1 2 Gibney, Elizabeth (23 January 2025). "China's cheap, open AI model DeepSeek thrills scientists". Nature. 638 (8049): 13–14. Bibcode:2025Natur.638...13G. doi:10.1038/d41586-025-00229-6. PMID 39849139. Archived from the original on 29 January 2025. Retrieved 12 February 2025.
  11. 1 2 Delbert, Caroline (31 January 2025). "DeepSeek Is Cracking the 'Black Box' of Corporate AI Wide Open". Popular Mechanics. Archived from the original on 13 February 2025. Retrieved 12 February 2025.
  12. Chen, Caiwei (12 February 2026). "What's next for Chinese open-source AI". MIT Technology Review. Retrieved 12 April 2026.
  13. 1 2 3 4 Metz, Cade; Tobin, Meaghan (23 January 2025). "How Chinese A.I. Start-Up DeepSeek Is Competing With Silicon Valley Giants". The New York Times. ISSN 0362-4331. Archived from the original on 23 January 2025. Retrieved 27 January 2025.
  14. 1 2 3 Metz, Cade (27 January 2025). "What is DeepSeek? And How Is It Upending A.I.?". The New York Times. ISSN 0362-4331. Archived from the original on 27 January 2025. Retrieved 27 January 2025.
  15. 1 2 "How China's PLA is gearing up to use DeepSeek AI models in actual combat". South China Morning Post. 23 March 2025. Retrieved 14 August 2026.
  16. Chen, Caiwei (24 January 2025). "How a top Chinese AI model overcame US sanctions". MIT Technology Review. Archived from the original on 25 January 2025. Retrieved 25 January 2025.
  17. Olcott, Eleanor; Wu, Zijing (24 January 2025). "How small Chinese AI start-up DeepSeek shocked Silicon Valley". Financial Times. Archived from the original on 25 January 2025. Retrieved 31 January 2025.
  18. 1 2 3 Ottinger, Lily (9 December 2024). "Deepseek: From Hedge Fund to Frontier Model Maker". ChinaTalk. Archived from the original on 28 December 2024. Retrieved 28 December 2024.
  19. Leswing, Kif (23 February 2023). "Meet the $10,000 Nvidia chip powering the race for A.I." CNBC. Archived from the original on 29 January 2025. Retrieved 30 January 2025.
  20. 独家|幻方量化回应市场关注:AGI不是用来炒股的,"和金融没关系". Yicai. Retrieved 3 February 2025.
  21. Yu, Xu (17 April 2023). "[Exclusive] Chinese Quant Hedge Fund High-Flyer Won't Use AGI to Trade Stocks, MD Says". Yicai Global. Archived from the original on 31 December 2023. Retrieved 28 December 2024.
  22. 1 2 3 Jiang, Ben; Perezi, Bien (1 January 2025). "Meet DeepSeek: the Chinese start-up that is changing how AI models are trained". South China Morning Post. Archived from the original on 22 January 2025. Retrieved 1 January 2025.
  23. McMorrow, Ryan; Olcott, Eleanor (9 June 2024). "The Chinese quant fund-turned-AI pioneer". Financial Times. Archived from the original on 17 July 2024. Retrieved 28 December 2024.
  24. 大模型价格又砍一刀 这次"屠夫"竟是量化私募?. www.cls.cn. 10 May 2024. Archived from the original on 27 December 2024. Retrieved 3 February 2025.
  25. Metz, Cade (23 February 2026). "Anthropic Accuses 3 Chinese Companies of Harvesting Its Data". The New York Times. ISSN 0362-4331. Retrieved 24 February 2026.
  26. Sultan, Abu; Kachwala, Zaheer (18 April 2026). Anil D'Silva; Devika Syamnath (eds.). "China's DeepSeek is raising funds at $10 billion valuation, The Information reports". Reuters. Retrieved 21 April 2026.
  27. "DeepSeek weighs new fundraising a month after closing first round". Financial Times. 14 July 2026. Retrieved 1 September 2026.
  28. Fan, Haze (14 July 2026). "DeepSeek Is Preparing For IPO Filing As Soon As This Year". Bloomberg. Retrieved 14 July 2026.
  29. "China's DeepSeek eyes US$70b valuation in new talks weeks after first round". South China Morning Post. 15 July 2026. Retrieved 26 August 2026.
  30. Forlini, Emily (30 January 2025). "DeepSeek Is Here to Stay as Microsoft, Perplexity Integrate Its Model". PCMAG. Retrieved 31 August 2026.
  31. Schneider, Jordan (27 November 2024). "Deepseek: The Quiet Giant Leading China's AI Race". ChinaTalk. Archived from the original on 29 November 2024. Retrieved 28 December 2024.
  32. "Despite China's 996 culture, DeepSeek founder Liang Wenfeng says his workers don't do overtime or even have KPIs: 'No one manages them'". Fortune.
  33. Mickle, Tripp; Swanson, Ana; Tobin, Meaghan; Metz, Cade (16 April 2025). "US Officials Target Nvidia and DeepSeek Amid Fears of China's A.I. Progress". The New York Times. ISSN 0362-4331. Retrieved 17 April 2025.{{cite news}}: CS1 maint: deprecated archival service (link)
  34. Greenspan, Anna; Konior, Bogna (2025). "Introduction: Fleeting Forces and Clever Machinations". In Bratton, Benjamin; Greenspan, Anna; Ireland, Amy; Konior, Bogna (eds.). Machine Decision is Not Final: China and the History and Future of Artificial Intelligence. Urbanomic, MIT Press. ISBN 9781913029999.
  35. Rai, Saritha, Loni Prinsloo, and Helen Nyambura "China's DeepSeek Is Beating Out OpenAI and Google in Africa" Bloomberg Technology. Accessed 27 October 2025.
  36. 1 2 An, Wei; Bi, Xiao; Chen, Guanting; Chen, Shanhuang; Deng, Chengqi; Ding, Honghui; Dong, Kai; Du, Qiushi; Gao, Wenjun; Guan, Kang; Guo, Jianzhong; Guo, Yongqiang; Fu, Zhe; He, Ying; Huang, Panpan (17 November 2024). "Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning". SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE. pp. 1–23. arXiv:2408.14158. doi:10.1109/SC41406.2024.00089. ISBN 979-8-3503-5291-7.
  37. 幻方力量 | 高速文件系统 3FS. High-Flyer. 13 June 2019. Archived from the original on 3 February 2025. Retrieved 3 February 2025.
  38. deepseek-ai/3FS, DeepSeek, 28 February 2025, archived from the original on 28 February 2025, retrieved 28 February 2025
  39. 1 2 3 hfreduce | 高性能的多卡并行通信工具. High-Flyer. 4 March 2020. Archived from the original on 28 January 2025. Retrieved 3 February 2025.
  40. "HFAiLab/hai-platform", High-Flyer, 2 February 2025, retrieved 3 February 2025
  41. 1 2 3 "DeepSeek Harness". GitHub. DeepSeek AI. Retrieved 29 August 2026. DeepSeek Harness (`dsh`) is an open-source agent harness ... It is built on an everything-is-a-plugin architecture and powered by Cordis
  42. 1 2 Lardinois, Frederic (13 August 2026). "DeepSeek open sources an agent harness where everything is a plugin". The New Stack. Retrieved 29 August 2026.
  43. "deepseek-ai/DeepSeek-V3.1 · Hugging Face". huggingface.co. 21 August 2025. Retrieved 25 August 2025.
  44. "DeepSeek-V3.1 Release | DeepSeek API Docs". api-docs.deepseek.com. Retrieved 25 August 2025.
  45. "deepseek-ai/DeepSeek-V3.1-Terminus · Hugging Face". huggingface.co. 22 September 2025. Retrieved 24 September 2025.
  46. Yuan, Jingyang; Gao, Huazuo; Dai, Damai; Luo, Junyu; Zhao, Liang; Zhang, Zhengyan; Xie, Zhenda; Wei, Y. X.; Wang, Lean (27 February 2025), Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, arXiv:2502.11089
  47. "deepseek-ai/DeepSeek-V3.2-Exp · Hugging Face". huggingface.co. 29 September 2025. Retrieved 2 October 2025.
  48. 1 2 Binder, Matt (3 December 2025). "DeepSeek v3.2: What it is, how it compares to ChatGPT, how to try it". Mashable. Retrieved 12 April 2026.
  49. 1 2 "DeepSeek-V3.2 Release". DeepSeek API Docs. 1 December 2025. Retrieved 12 April 2026.
  50. Field, Hayden (27 January 2025). "China's DeepSeek AI dethrones ChatGPT on App Store: Here's what you should know". CNBC. Archived from the original on 28 January 2025. Retrieved 27 January 2025.
  51. Picchi, Aimee (27 January 2025). "What is DeepSeek, and why is it causing Nvidia and other stocks to slump?". CBS News. Archived from the original on 29 January 2025. Retrieved 27 January 2025.
  52. 1 2 "LICENSE · deepseek-ai/DeepSeek-R1-0528". 28 May 2025. Retrieved 12 April 2026 via Hugging Face.
  53. 1 2 Sankaran, Vishwam (24 April 2026). "China's DeepSeek releases new AI model it claims beats all open-source competitors". The Independent. Retrieved 24 April 2026.
  54. 1 2 Chen, Caiwei (24 April 2026). "Three reasons why DeepSeek's new model matters". MIT Technology Review. Retrieved 25 April 2026.
  55. 1 2 Tolomia, Cris (24 April 2026). "DeepSeek is back with a new open-source AI model built on Chinese chips". Quartz. Retrieved 25 April 2026.
  56. "China's chipmakers rush to embrace DeepSeek's V4. Which names stand out?". South China Morning Post. 6 May 2026. Retrieved 7 May 2026.
  57. 1 2 Baptista, Eduardo (3 August 2026). "DeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says". Reuters. Retrieved 13 August 2026.
  58. 1 2 "Chinese AI startup DeepSeek quietly released 'DeepSeek V4 Pro 0813'". Gigazine. 13 August 2026. Retrieved 13 August 2026.
  59. 1 2 Lee, Chong Ming (10 September 2026). "DeepSeek says new Flash AI model beats Kimi K3 on cyber, coding benchmarks". South China Morning Post. Retrieved 10 September 2026.
  60. 1 2 Haralu, Kileuna (10 September 2026). "DeepSeek Launches V4.1-Flash to Make Large-Scale AI Inference Cheaper". Analytics India Mag. Retrieved 10 September 2026.
  61. "LICENSE · deepseek-ai/deepseek-coder-33b-base". Hugging Face. 28 October 2023. Retrieved 12 April 2026.
  62. "DeepSeek-MoE/LICENSE-MODEL". 11 January 2024. Retrieved 12 April 2026 via GitHub.
  63. "LICENSE · deepseek-ai/deepseek-math-7b-base". Hugging Face. 6 February 2024. Retrieved 12 April 2026.
  64. "LICENSE · deepseek-ai/deepseek-math-7b-instruct". Hugging Face. 6 February 2024. Retrieved 12 April 2026.
  65. "LICENSE · deepseek-ai/deepseek-math-7b-rl". Hugging Face. 6 February 2024. Retrieved 12 April 2026.
  66. "LICENSE · deepseek-ai/DeepSeek-V2.5". 5 September 2024. Retrieved 12 April 2026 via Hugging Face.
  67. "LICENSE-MODEL · deepseek-ai/DeepSeek-V3-Base". 26 December 2024. Retrieved 12 April 2026 via Hugging Face.
  68. "DeepSeek-Prover-V2/LICENSE-MODEL". 30 April 2025. Retrieved 12 April 2026 via GitHub.
  69. "deepseek-ai/deepseek-vl2". 27 November 2025. Retrieved 12 April 2026 via Hugging Face.
  70. <ref>"LICENSE · deepseek-ai/DeepSeek-R1-Distill-Qwen-32B". 20 January 2025. Retrieved 12 April 2026 via Hugging Face.
  71. "LICENSE · deepseek-ai/DeepSeek-V3.1-Base". 19 August 2025. Retrieved 12 April 2026 via Hugging Face.
  72. "LICENSE · deepseek-ai/DeepSeek-V3.1-Terminus". 22 September 2025. Retrieved 12 April 2026 via Hugging Face.
  73. "LICENSE · deepseek-ai/DeepSeek-Math-V2". 27 November 2025. Retrieved 12 April 2026 via Hugging Face.
  74. "LICENSE · deepseek-ai/DeepSeek-V3.2". 1 December 2025. Retrieved 12 April 2026 via Hugging Face.
  75. "README.md · deepseek-ai/DeepSeek-V4.1-Flash". 10 September 2026. Retrieved 10 September 2026 via Hugging Face.
  76. "Cyber and Artificial Intelligence Provisions in the FY2026 National Defense Authorization Act (NDAA)". www.congress.gov. Retrieved 6 July 2026.
  77. Sen. Cornyn, John [R-TX (18 December 2025). "Text - S.1071 - 119th Congress (2025-2026): National Defense Authorization Act for Fiscal Year 2026". www.congress.gov. Retrieved 6 July 2026.
  78. "PSPF Direction Update – DeepSeek Products, Applications and Web Services". protectivesecurity.gov.au/. 4 February 2025. Archived from the original on 7 February 2025. Retrieved 9 July 2026.
[edit]