Governments and major AI developers are increasingly treating who can use advanced models as a governance problem as much as a technical one. Over the past two years regulators and national security agencies have pushed model developers to adopt identity-, credential- and risk-based controls that restrict certain high-capability model behaviors to verified users or vetted organizations.
At the same time, commercial providers and hyperscalers are racing to redesign the hardware stack that powers inference, building bespoke accelerators and partnering with chipmakers to cut cost, latency and energy per query. That simultaneous move toward access control on the policy side and custom silicon on the infrastructure side is reshaping how AI will be governed and delivered at scale.
Why governments are vetting model users
Policymakers now see advanced generative models as dual‑use technologies that pose both broad economic opportunity and concentrated national‑security and cyber risks. High‑capability models can accelerate legitimate work, from code review to threat hunting, but they can also enable scalable attacks, disinformation campaigns and the automated design of exploits, prompting governments to demand stricter controls on access.
Regulatory frameworks have reinforced that view: the European AI Act and its associated guidance emphasize traceability, documentation and controlled access for higher‑risk systems, while U.S. agencies and guidelines (including NIST and federal executive directives) recommend staged releases and user verification for frontier models. These frameworks make vetting an expected governance tool rather than an ad hoc corporate decision.
National security actors have also exerted direct pressure. In several recent cases, governments have requested or required companies to limit distribution of particular model variants or to route access through vetted partner programs, moves that accelerated the adoption of identity‑based access programs by major providers. The practical result: some powerful model capabilities are now reachable only after an eligibility review or through specially authorized programs.
How developers implement user vetting in practice
Companies have adopted layered approaches that combine identity verification, organizational attestation, contractual obligations and runtime monitoring. OpenAI’s Trusted Access for Cyber (TAC) is an explicit example: it ties more permissive cyber‑capable model behavior to verified individuals and teams, requires stronger account security, and integrates programmatic safeguards.
Beyond identity checks, developers use staged releases and capability tiers: public, commercial, and trusted tiers with progressively stricter verification and logging. These staged models let vendors deliver defensive and research value to vetted users while reducing the surface area for widespread misuse. Industry commitments on frontier model safety have formalized staged releases into common practice.
Runtime telemetry and suspicious‑activity detection are also central. Providers route high‑risk queries through additional monitoring and apply dynamic restrictions when activity patterns suggest abuse, theft, or attempts to exfiltrate sensitive capabilities, a mix of automated classifiers and human review that supports revocation and post‑incident response.
The defensive case: widening defender access while constraining adversaries
One principal justification for vetting is to give legitimate defenders earlier and safer access to advanced capabilities. OpenAI and others have argued that identity‑based programs can accelerate patching, vulnerability discovery, and incident response by trusted teams while keeping the same capabilities out of the hands of malicious actors.
Vetted access programs are designed to scale: providers report expanding thousands of verified individuals and hundreds of teams into cyber‑focused tiers, enabling broader defensive use without wholesale public release of more permissive model behaviors. For defenders this creates a practical pathway to use advanced tooling under accountability structures.
However, these programs raise equity and operational questions. Smaller security teams, independent researchers and international partners may face higher barriers to entry, and governments worry about consistency in cross‑border vetting standards, all of which complicate the goal of democratizing defensive capability while managing risk.
Why firms are rushing to build custom inference chips
As model deployments scale, the economics of inference have become a dominant cost center for providers. Hyperscalers and leading model makers are therefore investing in custom inference silicon to reduce per‑query energy cost, lower latency, and gain supply‑chain leverage over a market long dominated by a single GPU vendor. Custom ASICs and reconfigurable accelerators promise both operational savings and architectural features tuned to modern LLM workloads.
Recent, high‑profile examples illustrate the trend: OpenAI’s Jalapeño (co‑developed with Broadcom) was unveiled in June 2026 as an inference‑optimized processor intended to run LLMs more efficiently at scale; hyperscalers such as Amazon and Google have similarly deployed generations of in‑house accelerators (Trainium and TPU/Trillium) to push down costs and lock in performance advantages.
Startups and alternative architectures also compete for inference economics. Companies such as Cerebras, SambaNova, Groq and Graphcore pitch novel dataflow or wafer‑scale designs that can outperform general‑purpose GPUs on certain inference patterns, and many of these vendors report commercial traction or strategic partnerships in 2025,2026. The result is a fractured but fast‑moving market at the inference layer.
Operational and market consequences of custom silicon
Custom chips shift bargaining power and can shorten or lengthen procurement cycles, depending on a buyer’s size and leverage. Large cloud providers can amortize design and NRE costs across massive fleets, but smaller customers may be pushed to vendor clouds or to commodity GPU markets if compatibility and availability diverge. That dynamic accelerates consolidation at the service layer while fragmenting the hardware ecosystem.
Supply constraints and manufacturing timelines matter. Even where a provider designs an ASIC, production depends on foundry capacity (notably TSMC) and packaging partners, which means multi‑year roadmaps and staged rollouts. Early generations often appear as mixed fleets, bespoke accelerators for in‑house workloads plus third‑party GPUs for external customers, rather than immediate wholesale replacement.
On performance and cost, bespoke inference silicon can deliver substantial gains per watt and per dollar for steady, high‑volume serving, but those benefits require software co‑design, tooling maturity and customer migration. The market therefore rewards providers who control both stack and scale, intensifying the incentive to vertically integrate compute, networking and model service layers.
Policy tensions and international implications
Custom chips and user vetting intersect with export controls, industrial policy and cross‑border competition. Governments that restrict model access or limit chip exports will shape not only who uses models but where and on what hardware those models run; the combination of access controls and sovereign silicon programs raises new questions about interoperability, safe‑harbor provisions and extraterritorial enforcement.
National security reviews (CFIUS‑style scrutiny) and export regimes complicate partnerships and sales; chip supply chains that cross borders can trigger regulatory second‑guessing of strategic partnerships. As a result, industry decisions about where to manufacture and whom to partner with now carry heightened geopolitical weight.
At the international level, regulators face a policy trade‑off between enabling vetted defensive uses and avoiding unintentional technology bifurcation that would hinder legitimate cross‑border collaboration. Harmonized standards for vetting, audit trails and reciprocal recognition would ease friction, but such cooperation is politically and technically difficult to achieve quickly.
What this means for enterprises, researchers and policymakers
For enterprises, the combination of user vetting and bespoke inference hardware means procurement strategies must consider both access policies and long‑term compute roadmaps. Organizations will weigh the benefits of lower per‑query cost against the operational friction of committing to a provider’s vetted programs and custom stack.
For researchers and smaller teams, vetted programs create a tighter gate for certain high‑capability features, shifting some work to open‑source stacks, private deployments, or partnerships with accredited institutions. Policymakers should therefore calibrate vetting so it minimizes malicious use while preserving research and defensive capacity for credible actors.
For policymakers, the twin trends underline a central choice: rely primarily on access‑control and vetting to manage risk, or couple those measures with market and supply‑chain interventions to limit hardware concentration and ensure resilience. Pragmatic approaches will combine clear vetting standards, auditability, and policies that keep critical infrastructure accessible to legitimate defenders.
In short, the near future of AI will be shaped as much by who is allowed to use models as by the silicon those models run on. The governance landscape is moving from abstract rules to concrete mechanisms, identity, contracts and chips, and the winners will be those who can translate policy constraints into operationally resilient platforms.
Policymakers, industry leaders and security practitioners should engage together to define interoperable vetting standards, promote responsible access for defenders, and ensure that vertical integration in silicon does not create brittle chokepoints. That coordination is the most practical path to harnessing powerful models while limiting their misuse.




