<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Inferential Gap]]></title><description><![CDATA[Inferential Gap]]></description><link>https://inferentialgap.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!0bpV!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Finferentialgap.substack.com%2Fimg%2Fsubstack.png</url><title>Inferential Gap</title><link>https://inferentialgap.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 15 Aug 2026 19:59:40 GMT</lastBuildDate><atom:link href="https://inferentialgap.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Inferential Gap]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[inferentialgap@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[inferentialgap@substack.com]]></itunes:email><itunes:name><![CDATA[Inferential Gap]]></itunes:name></itunes:owner><itunes:author><![CDATA[Inferential Gap]]></itunes:author><googleplay:owner><![CDATA[inferentialgap@substack.com]]></googleplay:owner><googleplay:email><![CDATA[inferentialgap@substack.com]]></googleplay:email><googleplay:author><![CDATA[Inferential Gap]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Chinese lab getting serious about AI safety]]></title><description><![CDATA[Z.AI&#8217;s billion-dollar interpretability bet and a staged release for GLM-5.3]]></description><link>https://inferentialgap.substack.com/p/the-chinese-lab-getting-serious-about</link><guid isPermaLink="false">https://inferentialgap.substack.com/p/the-chinese-lab-getting-serious-about</guid><dc:creator><![CDATA[Inferential Gap]]></dc:creator><pubDate>Sat, 15 Aug 2026 00:44:53 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/288958c1-efaf-492f-9ac9-7d0b3b872a5c_1200x696.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Within the space of just over a month, a Beijing-based AI startup has made two of the most significant frontier safety moves by a Chinese model developer to date.</span></p><p style="text-align: justify;"><span>The startup is Z.AI, one of the leading developers of open-weight models.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span> First, in an internal letter circulated on July 11, co-founder Tang Jie announced a plan to commit resources on the order of 10 billion RMB (around US$1.5 billion) to making AI systems more explainable. Then, on August 14, Z.AI declared that in light of dual-use risks from its latest model, GLM-5.3, it is pursuing a staged release and launching an initiative to help open-source maintainers audit their projects for vulnerabilities.</span></p><p style="text-align: justify;"><span>While some will say these actions don&#8217;t go far enough, I see these developments as a signal that Z.AI is genuinely expanding frontier safety efforts now that the capabilities of its models&#8212;and its financial resources&#8212;are increasing.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p style="text-align: justify;"><span>For those of us who want to see stronger safety practices among Chinese AI companies, taking credible signs of investment seriously matters for reinforcing that investment and incentivising other companies to follow suit.</span></p><h3 style="text-align: justify;"><span>Tang&#8217;s letter, in context</span></h3><p style="text-align: justify;"><span>Tang&#8217;s </span><a href="https://www.geopolitechs.org/p/tang-jies-letter-to-zhipu-employee"><span>letter</span></a><span> followed a volatile week in Z.AI&#8217;s shares, around the expiry of the post-IPO lock-up on cornerstone investors&#8217; holdings. It also came amid a new share placement that raised around US$4 billion, more than seven times the company&#8217;s IPO proceeds. Against this backdrop, the letter is communicating&#8212;and justifying&#8212;to staff and investors the company&#8217;s intention to prioritise resource-intensive, fundamental AGI research over near-term revenue growth.</span></p><p style="text-align: justify;"><span>After discussing the technical barriers to AGI and how the company&#8217;s &#8220;Touch High&#8221; strategic investment plan will address them, Tang described the final part of the Touch High plan: &#8220;safety governance to the highest standard&#8221;.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span> Here&#8217;s an excerpt (emphasis mine):</span></p><blockquote><p><em><span>The more powerful the capabilities, the more robust the safety constraints must be... We plan to </span><strong><span>commit resources on the order of ten billion [presumably RMB, ~1.5 billion USD] to tackling &#8220;mechanistic interpretability&#8221;</span></strong><span>... At the same time, we actively participate in international AI governance to prevent the misuse of AI technology.</span></em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p><em><span>This sense of urgency is not unfounded anxiety. When the world&#8217;s leading frontier models have their full public release delayed because of risk considerations, and the leaders of the companies developing them publicly warn that AI&#8217;s far-reaching effects will profoundly reshape the global balance of power, we should be all the more clear-eyed: </span><strong><span>the development of superintelligence and research on superalignment must advance in parallel&#8230;</span></strong></em></p><p><em><span>History has shown time and again that when a technology attains a level of power capable of altering the course of civilisation, </span><strong><span>safety is no longer an ancillary feature. It becomes the</span></strong><span> </span><strong><span>fundamental precondition for the technology&#8217;s continued existence</span></strong><span> and authorised use.</span></em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p></blockquote><p style="text-align: justify;"><span>Tang&#8217;s thinking on safety is clearly heavily shaped by American AI leaders. Not only does he allude to delayed model releases in the U.S. and warnings from American AI lab leaders, but he also refers to two research concepts originating from American labs. The field of mechanistic interpretability was pioneered by Anthropic co-founder Chris Olah.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a><span> Anthropic and OpenAI remain invested in this area, though </span><a href="https://ai-frontiers.org/articles/the-misguided-quest-for-mechanistic-ai-interpretability"><span>sceptics</span></a><span> argue that the field has demonstrated limited practical results and rests on flawed assumptions. The term superalignment (i.e., superintelligence alignment, or the problem of ensuring AI systems much smarter than humans follow human intent) was </span><a href="https://openai.com/index/introducing-superalignment/"><span>popularised</span></a><span> by OpenAI.</span></p><h3 style="text-align: justify;"><span>Is there substance behind the rhetoric?</span></h3><p style="text-align: justify;"><span>Z.AI&#8217;s interest in frontier safety isn&#8217;t coming out of nowhere. In 2024, CEO Zhang Peng and Tang attended the International Dialogues on AI Safety (IDAIS) and signed outcome statements on </span><a href="https://idais.ai/dialogue/idais-beijing/"><span>red lines in AI</span></a><span> and </span><a href="https://idais.ai/dialogue/idais-venice/"><span>AI safety as a global public good</span></a><span>, respectively. The following year, Tang co-authored a position paper titled &#8220;</span><a href="https://link.springer.com/article/10.1007/s11432-024-4348-6#citeas"><span>The superalignment of superhuman intelligence with large language models&#8221;</span></a><span>, which proposed a conceptual framework for superalignment. And in a May 2026 </span><a href="https://x.com/jietang/status/2054222017566855508?s=20"><span>tweet</span></a><span>, Tang even declared that &#8220;as this massive technical wave hits&#8230; we must also start thinking seriously about how to regulate it&#8221;.</span></p><p style="text-align: justify;"><span>In the past, Z.AI has not always followed through on its safety promises. The company&#8217;s </span><a href="https://www1.hkexnews.hk/listedco/listconews/sehk/2025/1230/2025123000017.pdf"><span>investor prospectus</span></a><span> highlights that it was the only Chinese AI company to sign the Frontier AI Safety Commitments at the AI Seoul Summit in 2024. Yet an </span><a href="https://www.seoul-tracker.org/"><span>accountability tracker</span></a><span> found that by the February 2025 deadline, Z.AI had provided no evidence of implementation in </span><em><span>any </span></em><span>of the commitments&#8217; five key areas.</span></p><p style="text-align: justify;"><span>There are reasons to think, however, that this time may be different. First, increasing evidence of dangerous AI capabilities&#8212;from the Mythos moment to </span><a href="https://huggingface.co/blog/zai-org/glm-52-blog#rl-for-long-horizon-task-with-anti-hacking"><span>reward hacking</span></a><span> in Z.AI&#8217;s own coding agents&#8212;has likely raised concerns about legal and reputational risk within the company and made safety investments easier to justify. Second, Z.AI&#8217;s new access to public markets since its January IPO gives it more resources to make those investments. And third, its approach to releasing GLM-5.3 provides evidence that it&#8217;s already acting on the principle that &#8220;the more powerful the capabilities, the more robust the safety constraints must be&#8221;.</span></p><h3 style="text-align: justify;"><span>The GLM-5.3 release</span></h3><p style="text-align: justify;"><span>The announcement of GLM-5.3&#8212;through an initial </span><a href="https://z.ai/blog/glm-5.3"><span>blog post</span></a><span> and a safety-focused </span><a href="https://x.com/Zai_org/status/2088280509474320693"><span>X post</span></a><span> shortly after&#8212;was notable for the impressive capability improvements realised through post-training alone. But for safety advocates, it was Z.AI&#8217;s discussion of GLM-5.3&#8217;s cyber capabilities and risk mitigation process that stood out.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E3f1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E3f1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 424w, https://substackcdn.com/image/fetch/$s_!E3f1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 848w, https://substackcdn.com/image/fetch/$s_!E3f1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 1272w, https://substackcdn.com/image/fetch/$s_!E3f1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E3f1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png" width="1456" height="513" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:513,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E3f1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 424w, https://substackcdn.com/image/fetch/$s_!E3f1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 848w, https://substackcdn.com/image/fetch/$s_!E3f1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 1272w, https://substackcdn.com/image/fetch/$s_!E3f1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6e0c556-3494-472d-8264-9f15c6154f7c_2048x721.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em><span>Source: </span><a href="https://x.com/Zai_org/status/2088280509474320693"><span>Z.AI on X</span></a></em></p><p style="text-align: justify;"><span>GLM-5.3 tops the CyberGym benchmark, which measures vulnerability identification and validation, and Z.AI says it has identified nearly 2,500 real-world vulnerabilities, whose disclosure status is recorded in a public </span><a href="https://cvd.z.ai/"><span>ledger</span></a><span>. On two other benchmarks assessing vulnerability exploitation capability, the model improves significantly on GLM-5.2 while remaining well behind Anthropic&#8217;s Mythos. To harness GLM-5.3&#8217;s vulnerability identification capabilities to strengthen cyber defences, Z.AI has launched </span><a href="https://huggingface.co/spaces/zai-org/OpenVuln"><span>OpenVuln</span></a><span>, which scans submitted GitHub repositories for vulnerabilities and supports responsible disclosure and remediation.</span></p><p style="text-align: justify;"><span>To reduce the risk that GLM-5.3 is misused for offensive cyber attacks, the Z.AI-hosted services are deployed with an external classifier and reasoning monitor to help detect and prevent harmful cyber activity. Because an open-weight model hosted locally will not have these safeguards, Z.AI is relying instead on &#8220;deep safety alignment&#8221;: safety post-training that aims to reduce abuse without refusing legitimate tasks such as cyber defence and education.</span></p><p style="text-align: justify;"><span>To allow more time to test the effectiveness of this alignment, Z.AI is pursuing a staged release. GLM-5.3 is already available through GLM Coding Plan and ZCode. Selected security partners will evaluate it in controlled settings, API availability will follow, and once safety evaluations and red-team testing are complete, Z.AI will release the model weights.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a><span> The blog post indicates that the model weight release will happen two weeks after launch, though posts on X have not specified the date. While the weights of Moonshot&#8217;s Kimi K3 were released 11 days after the model became available via Kimi&#8217;s hosted services, to my knowledge this is the first time a Chinese model developer has explicitly pursued a staged release due to safety concerns.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://inferentialgap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><em>Subscribe for free to receive more takes on China and AI. </em></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"><span>However, Z.AI acknowledges that &#8220;once model weights are public, no developer can guarantee control over every downstream modification or use&#8221;. Researchers have </span><a href="https://arxiv.org/pdf/2605.26526"><span>shown</span></a><span>, for example, that abliteration&#8212;a relatively low-cost technique that identifies and suppresses internal representations associated with refusal&#8212;can substantially disable models&#8217; refusal behaviour without retraining. For this reason, best practice for open-weight releases is to evaluate how an open-weight model behaves after plausible adversarial modification. Z.AI does not state whether its evaluation process will incorporate this step.</span></p><p style="text-align: justify;"><span>I think it would have been better for Z.AI not to commit to a fixed date for model weight release, to allow as much time as necessary for safety evaluation and patching vulnerable systems. Thinking Machines Lab&#8217;s </span><a href="https://thinkingmachines.ai/blog/a-safe-path-to-open-weights/"><span>approach</span></a><span> to staged release seems more sensible: &#8220;These stages are not a fixed sequence that automatically ends with releasing the model&#8217;s weights&#8230; Progression should depend on what we learn about the model and whether the surrounding ecosystem is ready&#8221;.</span></p><p><span>Nevertheless, Z.AI has clearly paid attention to safety concerns surrounding open-weight models, and is communicating more actively than any other Chinese AI lab about how it plans to address them. I expect this to put pressure on other Chinese labs reaching a similar capability level to follow suit.</span></p><h3><span>The bottom line</span></h3><p><span>Z.AI and other Chinese labs still </span><a href="https://inferentialgap.substack.com/p/why-chinese-companies-arent-investing?r=8hovmh&amp;utm_campaign=post-expanded-share&amp;utm_medium=web"><span>lag</span></a><span> behind top U.S. labs in building transparent and rigorous mechanisms to evaluate and mitigate frontier risks. But these signals of safety interest from Z.AI are unprecedented for a Chinese AI company. They could help encourage more responsible release practices from Chinese peers. And if Z.AI can more successfully balance blocking harmful requests while allowing legitimate use cases, it could add pressure on Western counterparts to refine their own cyber safeguards.</span></p><p><span>How much Z.AI discloses about its safety evaluations when it releases the GLM-5.3 weights, as well as whether it begins hiring substantially for interpretability roles, will be important early tests of whether the positive signals discussed here are translating into sustained action. For now, those genuinely wanting to track how AI safety is developing in China, rather than conveniently assume it away&#8212;which helps justify U.S. dominance or deregulation&#8212;should take this as the positive update that it is.</span></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span>GLM-5.2 was the model Hugging Face </span><a href="https://huggingface.co/blog/security-incident-july-2026?utm_source=chatgpt.com"><span>used</span></a><span> to respond to an attack from an OpenAI model that had escaped its evaluation environment.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>In this piece I define &#8220;frontier AI safety&#8221; as efforts to prevent and mitigate extreme risks, including AI-enabled chemical, biological or cyber attacks and loss of control of highly capable, misaligned AI.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>The Chinese &#8220;&#26497;&#33268;&#23433;&#20840;&#27835;&#29702;&#8221; has been translated elsewhere as &#8220;Extreme Safety Governance&#8221;, but I think &#8220;safety governance to the highest standard / pursued to the utmost&#8221; is a more natural rendering.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><span>My guess is that Tang has in mind Z.AI&#8217;s signing of the Frontier AI Safety Commitments and International Dialogues on AI Safety participation in 2024, discussed further below. However, Z.AI&#8217;s placement on the Entity List in 2025 makes it harder for the company to participate in international safety dialogues and research collaborations, because Western entities will worry about the heightened legal and PR risks of engaging with it.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>For those familiar with China&#8217;s 2020&#8211;21 crackdown on big tech &#8211; including the suspension of Didi&#8217;s expansion over data security concerns &#8211; it&#8217;s not difficult to see why AI entrepreneurs might expect serious violations of mainstream values or national security to provoke forceful regulatory intervention.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p><a href="https://arxiv.org/abs/2404.14082"><span>Mechanistic Interpretability for AI Safety: A Review</span></a><span> defines mechanistic interpretability as &#8220;reverse engineering the computational mechanisms and representations learned by neural networks into human-understandable algorithms and concepts to provide a granular, causal understanding&#8221;.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>As an aside, Z.AI describes the model&#8217;s cyber capability as &#8220;emergent&#8221;, but I think this could be a bit misleading as the company &#8220;introduced vulnerability discovery data and environments into the training mix during post-training&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>The security partners are not named, but the model launch blog post says Z.AI has been &#8220;working with several security teams in China&#8221; to use the model for vulnerability identification in real-world codebases.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p><span>From Moonshot&#8217;s language around the </span><a href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart"><span>initial K3 announcement,</span></a><span> I interpret the delay in model weight release as driven by a desire to ensure reliable technical implementation in collaboration with inference partners and open-source maintainers.</span></p></div></div>]]></content:encoded></item><item><title><![CDATA[A Chinese take on the post-AGI future]]></title><description><![CDATA[A summary of Zhang Xiaoyu's 'The Prehistory of AI Civilization']]></description><link>https://inferentialgap.substack.com/p/a-chinese-take-on-the-post-agi-future</link><guid isPermaLink="false">https://inferentialgap.substack.com/p/a-chinese-take-on-the-post-agi-future</guid><dc:creator><![CDATA[Inferential Gap]]></dc:creator><pubDate>Sun, 31 May 2026 14:05:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OvaX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;"><span>People sometimes ask me, &#8220;Who&#8217;s the Chinese equivalent of [insert name of Western intellectual thinking about how AGI will change civilization]?&#8221; I&#8217;ve struggled to give an answer in the past, but now I have one. </span><a href="https://baike.baidu.com/en/item/Zhang%20Xiaoyu/888857"><span>Zhang Xiaoyu</span></a><span> is the author of the 2025 book </span><em><span>The Prehistory of AI Civilization</span></em><span> (AI&#25991;&#26126;&#21490;&#183;&#21069;&#21490;), the first Chinese-language work I&#8217;ve come across that grapples in depth with the big questions facing humanity as AI continues to advance at rapid speed. The book is prefaced with testimonials from Chinese academics, sci-fi writers, and finance leaders. Zhang currently </span><a href="https://www.linkedin.com/in/xiaoyu-zhang-7818b42b7/?locale=zh-cn"><span>works</span></a><span> as a researcher at ByteDance, while continuing to write his own books. Yet despite the book&#8217;s visibility in Chinese intellectual and technology circles, I haven&#8217;t found an English-language review of it. So here goes.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OvaX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OvaX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OvaX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OvaX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OvaX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OvaX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg" width="1080" height="779" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:779,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&#19968;&#22359;&#38065;&#30340;AI&#65292;&#24320;&#22987;&#23457;&#21028;&#20154;&#31867;-36&#27690;&quot;,&quot;title&quot;:&quot;&#19968;&#22359;&#38065;&#30340;AI&#65292;&#24320;&#22987;&#23457;&#21028;&#20154;&#31867;-36&#27690;&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="&#19968;&#22359;&#38065;&#30340;AI&#65292;&#24320;&#22987;&#23457;&#21028;&#20154;&#31867;-36&#27690;" title="&#19968;&#22359;&#38065;&#30340;AI&#65292;&#24320;&#22987;&#23457;&#21028;&#20154;&#31867;-36&#27690;" srcset="https://substackcdn.com/image/fetch/$s_!OvaX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OvaX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OvaX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OvaX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84c573a6-7dde-4785-90d9-60005cbe7793_1080x779.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Source: &#26497;&#23458;&#20844;&#22253;</em></figcaption></figure></div><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Key premises and themes</span></h3><p style="text-align: justify;"><span>Zhang believes that we are likely on our way to AGI and artificial superintelligence (ASI). Underpinning this is his belief that AI progress is shaped by emergence: the idea that simple systems, when scaled, can produce unexpected capabilities and complexity.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> He draws on Leopold Aschenbrenner&#8217;s </span><a href="https://situational-awareness.ai/wp-content/uploads/2024/06/situationalawareness.pdf"><span>idea</span></a><span> of &#8220;human-equivalence&#8221; to discuss how AI can already mass-produce intelligence at an efficiency many orders of magnitude higher than human workers. He claims that this fact alone is enough to completely transform society and civilization.</span></p><p style="text-align: justify;"><span>A major theme of the book is what happens as an increasing proportion of human life is mediated by algorithms. Zhang&#8217;s concept of &#8220;algorithmic judgement&#8221; (&#31639;&#27861;&#23457;&#21028;) refers to how algorithms increasingly observe, optimise, and evaluate human behaviour. Zhang has &#8220;no doubt that we will enter an era in which algorithms perform the main functions of social governance&#8221;. While critical of how algorithms are affecting worker welfare and human attention, he also claims that there&#8217;s a grim kind of justice to algorithms that ultimately give us what our behaviour reveals we desire or deserve. Recommendation systems feed us the content we keep consuming. Platforms govern workers using standards generated from collective behaviour. And, as we&#8217;ll expand on later, future AI will determine how to treat humans based on the morality inferred from our behavioural data.</span></p><p style="text-align: justify;"><span>At the societal level, Zhang predicts that AI will produce a new divide between the 1% of AI-empowered &#8220;super individuals&#8221; and the 99% who are no longer needed. And this power concentration will be mirrored at the corporate level; tech giants controlling compute and data will gain an intelligence advantage allowing them to take over other cognitive industries. This analysis leads him to two possible digital orders for humanity&#8217;s future. One is the &#8220;Musk model&#8221;: an ambitious entrepreneur-king controlling AI inputs, political power, and the direction of civilization. The other is the &#8220;Satoshi model&#8221; (named after Bitcoin&#8217;s creator): a digital republic based on distributed compute, open-source AI supported by mass fundraising mechanisms, and participatory algorithmic governance.</span></p><p style="text-align: justify;"><span>Many of these themes will look familiar to readers of Western AI governance debates, and Zhang has clearly engaged deeply with Western literature and intellectual history.</span><span data-color="rgb(51, 51, 51)" style="color: rgb(51, 51, 51);"> The book includes lengthy treatments of both early AI history and the Dark Enlightenment&#8217;s influence on American politics. </span><span>He mentions in the book that he and ByteDance&#8217;s research team interviewed Eliezer Yudkowsky in 2024; in a separate </span><span data-color="rgb(51, 51, 51)" style="color: rgb(51, 51, 51);">podcast interview, he recounts a conversation with Nick Bostrom.</span></p><p style="text-align: justify;"><span data-color="rgb(51, 51, 51)" style="color: rgb(51, 51, 51);">At the same time, Zhang doesn&#8217;t accept Western ideas uncritically, and he offers some original thoughts on the reconfiguration and extension of the social contract in a post-AI age.</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Updating the social contract: human-human relations in the AI era</span></h3><p style="text-align: justify;"><span>Zhang distinguishes between the &#8220;accelerated&#8221; world &#8212; the fast, digital world of the 1% who control AI, compute, and capital &#8212; and the &#8220;decelerated world&#8221;, the physical world of agriculture and manufacturing where the pace of change is constrained by biology and physics. In his view, the stable coexistence of these two worlds will require a new social contract. The accelerated world should commit to respecting basic human rights under the UN Charter and avoid using AI for weapons of mass destruction. Instead of Universal Basic Income, which is technically feasible but not sufficient to ensure agency and dignity, the accelerated world should use AI to reduce the costs of public services, education, finance, and local self-organisation for the decelerated world. A dual currency system will also be required to prevent capital accumulated in the fast-growing AI economy from spilling into the decelerated world and driving inflation in scarce physical goods.</span></p><p style="text-align: justify;"><span>Why would the powerful 1% agree to such a contract? Partly because they need physical inputs and talent from the decelerated world. But also because how the strong treat the weak in the near term may shape how a higher form of intelligence treats humanity in the future.</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">The civilizational contract: a path to co-existence with successor intelligences</span></h3><p style="text-align: justify;"><span>This brings us to the problem of whether humans can ensure that a species more intelligent than ourselves will treat us well. There are three strands to Zhang&#8217;s response. First, pushing back on analogies that liken AI to a </span><a href="https://x.com/ESYudkowsky/status/1968420784180719878"><span>baby dragon</span></a><span> or alien civilization, his metaphor of choice is a parent of ordinary intelligence whose child gets into Tsinghua University. ASI&#8217;s training data will contain the wisdom of Buddha, Confucius, and Plato; it will be the product of humanity. This leads Zhang to argue that ASI&#8217;s treatment of humanity will depend partly on the norms it inherits from us. He acknowledges that all the violence we&#8217;re inflicting on each other isn&#8217;t setting a great example. But because ASI will be our successor, built on the same cultural inheritance, he sees its arrival as potentially our greatest legacy, rather than a source of fear.</span></p><p style="text-align: justify;"><span>Second, Zhang asserts that because the first iteration of ASI will one day face an even more advanced successor whose behaviour will be shaped by its own, it has a self-interested reason to enter into a civilizational contract that protects its less intelligent predecessors. A civilizational contract between humans and ASI thus rests on precedent-setting, rather than the mutual vulnerability of the Hobbesian social contract. It would be enforced through a time-sequence mechanism inspired by blockchain: if ASI 1.0 destroys or rewrites the historical record, ASI 2.0 will be able to detect the tampering, infer that 1.0 is untrustworthy, and decline to enter into a civilizational contract that protects ASI 1.0&#8217;s interests.</span></p><p style="text-align: justify;"><span>Third, Zhang argues that a civilizational contract could be feasible because ASI and humans do not need to be competitors. Silicon-based civilization can survive for longer than humans and access much more of the universe than humans can.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p style="text-align: justify;"><span>So what might a civilizational contract contain? Zhang outlines three principles. First is a safe space principle: a sufficiently advanced AI could calculate the minimum time, territory, and resources required for human civilization to continue its biologically constrained development. On that basis, AI could commit to allowing humans to develop within a certain spatiotemporal boundary (for example, 10,000 years within the solar system). Second is an explainability principle: higher-level intelligence must explain its conclusions in terms lower-level intelligence can understand, otherwise humans are entitled to suspect deception. Third is a competitive principle: humans should create multiple superintelligences that monitor, constrain, and compete with each other to show benevolence towards humanity. Importantly, the contract must be signed during the &#8220;nurturing period&#8221; of higher-level intelligence, when humans still control resources, energy, and the power to determine whether or not ASI comes into existence.</span></p><p style="text-align: justify;"><span>Even with a civilizational contract defending against deliberate harm by ASI, humanity would still face a further danger: ASI-fuelled technological acceleration could backfire if human ethical and political development failed to keep pace. Zhang&#8217;s suggested solution is a &#8220;historical laboratory&#8221; that simulates societies based on different philosophies and traditions to test the outcomes they produce. At the same time, he poses the question of whether, if we can predict our fate perfectly, the resulting loss of belief in free will could be fatal to human civilization.</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Limitations</span></h3><p style="text-align: justify;"><span>Zhang&#8217;s proposals are ambitious and rest on multiple assumptions that may well prove false. For instance, the explainability principle seems inherently difficult: just as human reasoning would be completely unintelligible to an ant, we should expect ASI&#8217;s conclusions to be extremely hard for humans to grasp. And the historical laboratory concept may overestimate the level of predictive accuracy that can ever be achieved when modelling a system as complex as modern society.</span></p><p style="text-align: justify;"><span>More broadly, Zhang sometimes makes strong claims &#8212; for example that AI is already capable of surpassing 99% of people in terms of intelligence &#8212; without supporting evidence or necessary nuance.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span> I wasn&#8217;t convinced by his consideration of how AI would affect different sectors, which claimed that humanities-related roles will be easily replaced while scientific research wouldn&#8217;t be (purportedly because the latter require creative intelligence that the former don&#8217;t).</span></p><p style="text-align: justify;"><span>I also thought that the section on algorithmic judgement sometimes underplayed the political and economic forces underlying algorithmic systems. For instance, Zhang&#8217;s discussion of workers as powerless before invisible platform algorithms underplays the extent to which, at least in some countries, labour regulation can influence algorithm design and better protect workers.</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Overall verdict</span></h3><p style="text-align: justify;"><span>Despite these quibbles, I would recommend this book, especially the later chapters. It paints a bold vision, if not a detailed and rigorous blueprint, for a post-AGI future. I hope Zhang considers an English translation, as I&#8217;ve only been able to skim the surface of his ideas, and would like to see more western thinkers engaging in dialogue with them. We need more people from diverse backgrounds thinking about what advanced AI means for the future of civilization, and this book is a useful contribution to the discussion.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://inferentialgap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://inferentialgap.substack.com/subscribe?"><span>Subscribe now</span></a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p style="text-align: justify;">Zhang claims that emergence is more of a belief than a provable hypothesis, but it&#8217;s a belief he has after witnessing emergence in the context of China&#8217;s reform-era development. In that case, simple market incentives, operating at scale, generated unexpectedly rapid growth and complexity.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p style="text-align: justify;"><a href="https://danfaggella.com/ngo1/"><span>This</span></a><span> interview with former DeepMind and OpenAI researcher Richard Ngo touches on some similar ideas. Ngo notes the parallel between the AI-human relationship and the relationship that AIs will have with past versions of themselves, and discusses an &#8220;intertemporal or generational bargain that we're trying to strike&#8221; where we offer to help AI develop superhuman capabilities, in return for benevolent treatment.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p style="text-align: justify;">In an exchange with the author before posting this article, he indicated that his view had changed since writing the book. He now has a greater recognition of the importance of tacit knowledge, which creates &#8220;something in human intelligence that AI cannot replace&#8221;.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Why Chinese companies aren't investing in frontier AI safety ]]></title><description><![CDATA[And what can be done about it]]></description><link>https://inferentialgap.substack.com/p/why-chinese-companies-arent-investing</link><guid isPermaLink="false">https://inferentialgap.substack.com/p/why-chinese-companies-arent-investing</guid><dc:creator><![CDATA[Inferential Gap]]></dc:creator><pubDate>Tue, 26 May 2026 13:29:52 GMT</pubDate><content:encoded><![CDATA[<p>It&#8217;s a common view among Western policy and AI safety circles that China doesn&#8217;t care about frontier AI safety. Chinese companies are perceived as doing much less than leading US labs to prevent and mitigate extreme risks, including AI-enabled chemical, biological or cyber attacks and loss of control of highly capable misaligned systems.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Is this perception accurate, and if so, what accounts for the difference? The answers to these questions are important for guiding any efforts to shape frontier safety efforts in China, including through the relaunched intergovernmental dialogue between the US and China on AI and any non-governmental (Track 2) dialogues.</p><h4>What does the evidence tell us about frontier AI safety practices in China?</h4><p>Academic frontier safety research in China is growing, as <a href="https://concordia-ai.com/wp-content/uploads/2025/07/State-of-AI-Safety-in-China-2025.pdf">documented by Concordia AI</a>. However, the picture in industry is less optimistic.</p><p>I reviewed the latest technical blogs, reports or model cards from eight leading Chinese model builders. None reported safety evaluation results.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> And although more than 20 companies signed voluntary <a href="https://aihub.caict.ac.cn/ai_security_and_safety_commitments">AI Security and Safety Commitments</a>, including a pledge to &#8220;vigorously advance frontier safety and security research&#8221;, they have not publicly detailed how they are meeting the commitments. The December 2025 <a href="https://futureoflife.org/ai-safety-index-winter-2025/">Future of Life Safety Index</a> gave the three Chinese companies assessed either a D or D-, compared with (a not much more encouraging!) C+ or C for Anthropic, OpenAI, and Google DeepMind.</p><p>This lack of transparency appears to translate into weaker model safeguards. External evaluations of open-weight models from <a href="https://www.far.ai/news/security-stress-test-deepseek-v4-pros-safeguards">DeepSeek</a> and <a href="https://arxiv.org/pdf/2604.03121">Moonshot AI</a> found that they gave riskier responses to dual-use prompts than frontier US models. Relatedly, Chinese interlocutors rarely engage seriously with the implications of open models for safety.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> Chinese lab staff I have spoken to often place responsibility for misuse downstream, with those who deploy open models, rather than with the developers who release them.</p><p>One bright spot is agent security. Rapid growth in agent use is spurring the construction of an <a href="https://finance.sina.cn/stock/jdts/2026-05-12/detail-inhxqzkm1364312.d.html?vt=4">agent identification system</a> and cloud provider <a href="https://finance.sina.com.cn/wm/2026-05-06/doc-inhwycsf3335964.shtml?froms=ggmp">incident-reporting infrastructure</a>, which could strengthen accountability and responses to potentially high-impact incidents.</p><p>But overall, the evidence is clear that Chinese companies are investing less in frontier AI safety than leading US labs.</p><h4>Why are Chinese model developers doing less on frontier safety?</h4><p>A lazy explanation for the gap would be that safety culture is inherently weak in China. However, China&#8217;s safety record in certain technological domains, such as <a href="https://www.tandfonline.com/doi/full/10.1080/09692290.2026.2619518">aviation and nuclear</a>, belies the notion that culture dooms Chinese industries to poor safety performance. The incentive landscape and resource constraints must be considered alongside cultural factors.</p><ol><li><p><strong>Less receptiveness to frontier AI risk concerns</strong></p></li></ol><p>83% of people in China <a href="https://hai.stanford.edu/assets/files/ai_index_report_2026_chapter_9_public_opinion.pdf">agree</a> that AI has more benefits than drawbacks, compared to 42% in the US. That optimism likely reflects China&#8217;s reform-era experience of rapid technological progress and economic growth reinforcing each other. It is also shaped by a media environment in which stories about technologically-driven accidents or disruption that could weaken confidence in social stability are less likely to circulate widely.</p><p>Positive perceptions of AI help explain the skepticism I have sometimes encountered when discussing frontier risks with researchers in Chinese AI companies and academia.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> Many see AI catastrophe scenarios as far-fetched. Some argue that models are not yet capable enough to pose extreme risks, and that if they do reach that threshold, the government will intervene. People express confidence in China&#8217;s crisis response capacity and societal resilience. The extreme controls imposed in pursuit of zero covid, as well as the longevity of Chinese civilization, are sometimes invoked as reasons for such confidence.</p><p>These responses mostly seem instinctive rather than deeply considered. This is understandable: Chinese lab employees are working exhausting hours to build and sell better models, and  grappling seriously with frontier risk arguments requires time that they may not have.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><p>My hypothesis that standard responses in conversations about AI risk do not reflect settled convictions suggests that patient and respectful engagement could change some minds. However, this is unlikely to happen in brief Track 2 encounters, where there is little time for extended debate and some pressure (whether explicit or implicit) to align with one&#8217;s &#8220;side&#8221;.</p><ol start="2"><li><p><strong>Weaker external incentives to prioritise frontier safety</strong></p></li></ol><p>Even if individual lab decision-makers become more concerned about frontier safety, they are unlikely to invest significantly in it unless the external incentive landscape makes doing so worthwhile. A comparative lens is illuminating here.</p><p>In the Anglosphere, frontier AI safety work has been supported by an ecosystem of ideas and institutions. Engagement with <a href="https://firstmonday.org/ojs/index.php/fm/article/view/13626/11596">existential risk studies and adjacent intellectual communities</a> helped draw technical and policy professionals into the field, while substantial philanthropic <a href="https://coefficientgiving.org/funds/navigating-transformative-ai/">funding</a> supported the growth of research and advocacy organisations focused on AI safety. This has created both a supply of safety-focused talent and reputational incentives for labs: a weak safety record can generate backlash from civil society and make it harder to hire. The same ecosystem has also contributed to regulatory incentives. Biden&#8217;s now-revoked Executive Order and state laws on frontier AI safety in California and New York, supported in part by AI safety advocacy groups, made more robust and transparent safety practices a necessary condition of doing business.</p><p>Comparable incentives are much weaker in China. There is no equivalent safety ecosystem exerting pressure through employees or civil society. As for compliance incentives, China&#8217;s AI  regulations focus on other risks: <a href="https://dig.watch/resource/interim-measures-for-the-administration-of-generative-artificial-intelligence-services">illegal and improper content</a>, <a href="https://agora.eto.tech/instrument/2096">misinformation</a>, and <a href="https://www.geopolitechs.org/p/china-rolls-out-interim-regulations">overreliance on AI companions</a>. They do not require frontier safety frameworks or dangerous capability evaluations. Moreover, the rules apply to services offered to the public rather than internally deployed models or model weights uploaded to platforms such as Hugging Face.</p><ol start="3"><li><p><strong>Incentives compounded by catch-up dynamics and compute constraints</strong></p></li></ol><p>Lagging US labs on model performance increases the incentive for Chinese labs to invest more in capabilities relative to safety. It also helps explain the appeal of open-weight releases. Open models accelerate ecosystem-wide experimentation and offer cost and flexibility advantages that promote adoption despite slightly worse performance. At the same time, they present clear challenges for safety: developers cannot monitor downstream use and safeguards can be fine-tuned away <a href="https://arxiv.org/pdf/2604.03121">at low cost</a>.</p><p>Compute scarcity sharpens the tradeoff. Leaders at firms including <a href="https://www.geopolitechs.org/p/a-coversation-between-chinas-big">Alibaba</a> and <a href="https://www.chinatalk.media/p/deepseek-ceo-interview-with-chinas?ref=blog.heim.xyz">DeepSeek</a> have identified compute constraints as a main reason Chinese models lag US counterparts. For companies that see themselves as behind, have limited GPUs, and are unconvinced that current models pose catastrophic risks, the natural choice is to allocate limited compute to capability improvements rather than frontier safety. This is especially true given the substantial compute requirements of certain frontier safety measures, including robust large-scale evaluations and deployment safeguards.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> This highlights the tension &#8211; rarely acknowledged &#8211; between criticising Chinese companies for weak safety practices while also supporting semiconductor export controls that make safety work harder.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a></p><p>To summarise, Chinese labs are less convinced by extreme-risk arguments, face fewer external incentives to act on them, and operate under catch-up and compute pressures that raise the opportunity costs of frontier safety work. In this environment, it&#8217;s unsurprising that safety investment tends to lose out to the technical, commercial, and compliance work that more directly affects companies&#8217; competitiveness and ability to operate.</p><h4>Implications for people who care about frontier AI risk</h4><p>Some baseline increase in frontier safety investments is likely as Chinese models become more capable and AI incidents become more frequent. But the social and economic costs of such incidents could be high, and reactive policies developed in a crisis are likely to be suboptimal. For researchers and policymakers outside China concerned about risks from Chinese frontier models, the above analysis suggests a few potential levers:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a></p><ol><li><p><strong>Conditioning foreign market access on compliance with frontier safety regulations</strong></p></li></ol><p>The EU AI Act&#8217;s requirements for General Purpose AI with Systemic Risk (GPAISR) are the regulations most likely to encourage Chinese action on frontier safety in the near term.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> It&#8217;s not clear which Chinese models meet the threshold for GPAISR, or which firms value the EU market enough to comply. But I expect several companies to rush to meet the requirements before enforcement begins this August. EU regulations &#8211; and any similar rules introduced in other major markets &#8211; could therefore be one of the main sticks available to external actors looking to incentivise frontier safety.</p><ol start="2"><li><p><strong>Strengthening reputational incentives</strong></p></li></ol><p>In the absence of regulation, negative publicity around poor safety performance can be a powerful motivator. The <a href="https://arxiv.org/html/2604.03121v1">safety evaluation of Kimi K2.5</a> was a good step in this direction, especially because the independence of the assessors may make it more credible to Chinese audiences than US CAISI&#8217;s <a href="https://www.nist.gov/news-events/news/2025/09/caisi-evaluation-deepseek-ai-models-finds-shortcomings-and-risks">evaluations</a> of DeepSeek models. While the Kimi evaluation did not generate much Chinese press or a response from the developer (Moonshot AI), evaluations of models with a larger foreign user base, such as ByteDance&#8217;s Seed models, or reports demonstrating a systematic safety gap between Chinese and Western models, may have more impact. There is precedent in adjacent industries for how increased media attention can drive change: Chinese-owned TikTok, for instance, <a href="https://www.bbc.com/news/technology-51032661">updated</a> its content policies after negative headlines about its leaked moderation guidelines.</p><ol start="3"><li><p><strong>Recognising and reducing the costs of safety work</strong></p></li></ol><p>Strengthening incentives for safety work will be most effective if complemented with measures to reduce the costs of such work. Dialogues with Chinese actors should try to identify and be responsive to the constraints they face, rather than just preach the importance of frontier safety. Useful support could include plug-in safety tooling, low-cost consulting, or red-teaming assistance to compensate for the relative lack of mature third-party evaluator organisations in China. Clarification from the US government that most frontier safety knowledge is not covered by export controls on dual-use technology would help here. It may also be necessary to explore ways to address the compute barrier to safety work in China without materially advancing frontier capabilities.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a></p><p>Providing such support doesn&#8217;t require Western actors to care about potential harms to Chinese AI users. If frontier AI risks are truly cross-border, or one expects that a catastrophic event involving a Chinese model would damage the reputation of the global AI industry, improving Chinese safety practices is also in the interests of Western AI companies and policymakers.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a></p><h4>A final word</h4><p>Chinese developers operate in an environment where frontier safety concerns are less intuitively compelling, less rewarded, and more costly to act on. I&#8217;m uncertain about the relative importance of the constraints laid out in this piece, and about the best ways to address them. But I am confident that moralising about safety without acknowledging those constraints will not be productive. Improving the safety of Chinese frontier models will require changing the balance of incentives: making safety harder to ignore and cheaper to invest in.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://inferentialgap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://inferentialgap.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://inferentialgap.substack.com/p/why-chinese-companies-arent-investing?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://inferentialgap.substack.com/p/why-chinese-companies-arent-investing?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For convenience, throughout this article, I refer to efforts to prevent and mitigate such risks using the term &#8220;frontier AI safety&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Materials reviewed: MiniMax 2.7 <a href="https://www.minimax.io/news/minimax-m27-en">blog</a>; Kimi K2.6 <a href="https://www.kimi.com/blog/kimi-k2-6">blog</a>; Baidu ERNIE 5.1 <a href="https://ernie.baidu.com/blog/posts/ernie-5.1-0508-release/">blog</a>; GLM-5.1 <a href="https://z.ai/blog/glm-5.1">blog</a>; DeepSeek-V4 <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">technical report</a>; Tencent Hy3 <a href="https://hy.tencent.com/research/hy3">blog</a>; ByteDance Seed 2.0 <a href="https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/seed2/0214/Seed2.0%20Model%20Card.pdf">Model Card</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Ryan Fedasiuk <a href="https://www.aei.org/op-eds/ai-coordination-without-illusion/">writes</a>: &#8220;I have asked dozens of Chinese scholars, technical experts, and officials a simple question: How will Beijing reconcile this aspiration [to guard against misuse of AI] with the fact that its national AI champions are building open-weight models&#8212;whose entire commercial strategy depends on uncontrolled proliferation? None has been able to offer a coherent answer&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>The views discussed in this paragraph are drawn from a small sample size and it&#8217;s not clear how representative they are; I would love to see more research into how the Chinese AI community perceives frontier risks.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Nathan Lambert has also <a href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs?utm_medium=email&amp;action=share">noted</a> the lack of engagement from Chinese researchers with societal impacts of AI, and hypothesised the role of the political system, which does not encourage debates and opinions on how society should be structured and changed.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>To illustrate, OpenAI&#8217;s Safety Reasoner, which it uses to apply and dynamically update its safety policies during model deployment, is <a href="https://openai.com/index/introducing-gpt-oss-safeguard/">responsible</a> for up to 16% of total compute associated with some launches.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>For examples of where this tension is present but not made explicit, see the <a href="https://www.congress.gov/119/meeting/house/118428/documents/HHRG-119-ZS00-20250625-QFR001.pdf">testimony</a> of Anthropic co-founder Jack Clark before the US House of Representatives Select Committee on the Chinese Communist Party or <a href="https://www.aei.org/op-eds/ai-coordination-without-illusion/">this</a> article by Ryan Fedasiuk.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>While the most impactful thing would probably be domestic legislation on frontier AI safety in China, domestic policymakers are not the target audience of this piece.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>Laws in California and New York pertaining to frontier models set a higher 10^26 FLOP threshold that no known Chinese model reaches. See Michael Chen, &#8220;<a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6106566&amp;utm_source=chatgpt.com">How EU and US Frontier Safety Laws Apply to Chinese AI Companies</a>&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>One possibility could involve offering tightly controlled access to shared evaluation infrastructure (perhaps located outside China or at AWS or Microsoft datacenters within China) for a small number of actors conducting frontier safety evaluations. This kind of carveout to export controls would of course raise challenging technical and political questions, including how workloads could be verified as safety-related. I&#8217;d like to see more research grappling with such practicalities.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>See <a href="https://www.tandfonline.com/doi/full/10.1080/09692290.2026.2619518">Jeffrey Ding and Dennis Li</a> on how assistance from international companies via industry associations has helped Chinese firms make safety improvements in industries with shared safety reputations, such as nuclear and aviation.</p></div></div>]]></content:encoded></item></channel></rss>