Liang Wenfeng, founder of Chinese AI lab DeepSeek, opened 2026 the way he closed out 2025: with a technical research paper rather than a product launch or a public appearance. In January, DeepSeek published a paper co-authored by Liang proposing a new training architecture called Manifold-Constrained Hyper-Connections (mHC), aimed at making large language model training meaningfully more compute-efficient — part of a consistent pattern in which DeepSeek’s published research has served as an early signal of engineering choices that later show up in its flagship model releases, rather than pure academic output disconnected from the company’s product roadmap.
That pattern held again in April 2026, when DeepSeek released its next-generation flagship model, DeepSeek V4, in two variants: V4-Pro, a 1.6-trillion-parameter model built for what the company calls “Expert Mode,” and V4-Flash, a smaller 284-billion-parameter model for faster, lower-cost “Fast Mode” use. Both variants shipped with a 1-million-token context window, a substantial increase over DeepSeek’s prior generation of models, and both landed on the late-April timeline Liang had reportedly communicated internally months earlier — a release cadence that reinforced DeepSeek’s reputation, established after its January 2025 debut sent shockwaves through global AI markets, for shipping genuinely capable models on a leaner compute and financial footprint than most Western frontier labs.
A More Difficult Story Behind R2
DeepSeek R2, the anticipated successor to the R1 reasoning model that first established the company’s global reputation, followed a considerably rockier path. As of mid-2026, R2 had not been officially released, and reporting from Reuters and The Information indicated Liang personally held back the model because its performance did not meet his own bar for release — a notable instance of a founder publicly associated with rapid, confident shipping choosing restraint over hitting a widely anticipated launch window. Part of the delay traced to a specific technical setback: DeepSeek had attempted to train R2 using Huawei’s Ascend AI chips, part of a broader effort by Chinese AI labs to reduce dependence on Nvidia hardware amid continued U.S. export restrictions, but the training run reportedly failed, forcing DeepSeek to pivot back to Nvidia GPUs for training while continuing to use Ascend chips for the comparatively less demanding task of running inference on already-trained models.
A Founder Who Rarely Speaks Publicly
Liang has remained one of the most publicly reserved major AI lab founders globally, rarely granting interviews or making conference appearances even as DeepSeek’s models have repeatedly drawn global attention and, at points, moved markets. He founded DeepSeek in 2023 as an offshoot of High-Flyer, a Chinese quantitative hedge fund he had previously founded and built substantial computing infrastructure for — a background in quantitative finance and large-scale computing infrastructure that gave DeepSeek an unusual starting position compared with most AI labs built directly around AI research from the outset. That hedge-fund lineage has also been widely credited with shaping DeepSeek’s distinctive engineering culture, which has consistently emphasized training efficiency and cost discipline over simply scaling up compute spending to match better-funded Western competitors.
The Chip Question at the Center of China’s AI Strategy
DeepSeek’s reported Ascend training failure carries significance well beyond the company’s own roadmap: it is one of the more concrete, publicly reported data points suggesting that domestic Chinese AI chips still lag Nvidia hardware for the most demanding large-scale training workloads, even as Chinese chipmakers and the Chinese government have pushed hard to reduce reliance on U.S. semiconductor technology. DeepSeek’s pragmatic response — continuing to use Nvidia for training while shifting inference workloads to Ascend — reflects a hybrid approach likely to be echoed across other Chinese AI labs navigating the same tradeoff between chip availability, cost, and raw training performance in the near term.
A Model That Reshaped Global AI Investment
DeepSeek’s January 2025 release of its R1 reasoning model triggered a sharp, widely covered selloff in U.S. AI and semiconductor stocks, driven by investor concern that a Chinese lab had achieved near-frontier reasoning performance at a small fraction of the compute cost associated with leading U.S. models. That single event elevated Liang from a relatively obscure quantitative-finance figure to one of the most closely watched individuals in global AI within a matter of days, a level of scrutiny his own consistently low public profile has done little to diminish since.
What This Means for Philippine Founders
DeepSeek’s models remain among the most accessible and cost-effective large language models available to Philippine developers and startups building AI features on a limited budget, given DeepSeek’s open-weight release strategy and consistently lower per-token pricing compared with leading U.S. labs. Liang’s decision to delay R2 rather than ship a model that didn’t meet his own performance bar is also a useful, concrete example for Filipino AI founders under pressure to move fast: even one of the most closely watched labs in the world, backed by real technical credibility, chose to hold a flagship release back rather than compromise on quality — a tradeoff worth weighing seriously against the instinct to ship on a fixed timeline regardless of readiness.
Share this article