The AI Supply Chain Has a Supply Chain (Problem)
Why model provenance, infrastructure, and jurisdiction now belong in enterprise AI due diligence.There was a time when buying software was relatively simple. You bought the software, installed the software, and complained about the software.Then came the cloud, which meant you also needed to know…
Why model provenance, infrastructure, and jurisdiction now belong in enterprise AI due diligence.There was a time when buying software was relatively simple. You bought the software, installed the software, and complained about the software.Then came the cloud, which meant you also needed to know where the software ran, who hosted it, where the data went, who could see it, what subcontractors were involved, whether everyone had the right certifications, and whether some guy named Randy still had administrator access from a consulting engagement three years ago.Artificial intelligence has managed to make all of this even less simple.American companies evaluating AI vendors have learned to ask a familiar set of questions: Where is our data stored? Is it encrypted? Is our information used to train the model? Who can access it? Does the vendor have SOC 2 certification? Can we control retention? What happens to our data when the contract ends?Those are reasonable questions, but they increasingly cover only part of the problem. They concentrate on the vendor and the movement of customer data. AI introduces another layer below the vendor itself: the models, infrastructure, tools, libraries, and service providers that make the product work.For law firms, this is already more than a technology-procurement issue. A firm may hold a client’s merger plans, trade secrets, litigation strategy, details of a cyberattack, government contracting information, intellectual property and privileged internal communications. When that material enters the firm’s technology environment, the law firm becomes part of the client’s information supply chain, and the technology suppliers used by the firm become extensions of it.That makes a recent development in legal AI particularly interesting.Harvey, one of the most prominent companies in the field, recently announced Harvey Tenet, its first post-trained open-weight model for legal work. Harvey says Tenet improves materially on the underlying model in several legal benchmarks and is designed for long-horizon legal tasks. The more interesting detail, at least from a security, procurement, and governance perspective, is what sits underneath it.Harvey Tenet starts with Kimi K3.Kimi K3 was developed by Moonshot AI, a Beijing-based Chinese artificial intelligence company. Harvey did not simply call Kimi through an API. In its technical description, Harvey says that it and Fireworks AI post-trained the Kimi K3 base using reinforcement learning for long-horizon legal work. Kimi K3 is therefore a foundational component of the resulting Harvey model.There is no basis to suggest that Harvey is (necessarily) sending client information to China or that Moonshot has access to Harvey customers’ matters. An open-weight model can be downloaded, hosted, and operated independently of its original developer.The absence of data transmission to China, however, does not exhaust the procurement question. An American customer may still reasonably want to understand the origin, training, alignment, behavior and future availability of critical technology developed in a country that competes directly with the United States for leadership in advanced AI.That is where this becomes a much larger issue than legal technology.We Have Seen This BeforeAmerican cybersecurity policy has spent much of the last decade teaching companies to look below the name on the invoice. Cloud providers matter. Subcontractors, managed service providers, software libraries and embedded components matter too.SolarWinds made the lesson difficult to ignore. A trusted vendor can inherit risk from somewhere much deeper in the development and delivery chain. Excellent security controls at the customer and careful diligence on the vendor do not eliminate a vulnerability that entered upstream.AI makes the dependency chain unusually complicated. Behind a user-facing product there may be a cloud platform, an inference provider, open-source libraries, retrieval infrastructure, embedding models, orchestration software, vector databases, and one or more foundation models. Some of those components depend on still other providers.Each piece has an owner, origin, development history, security posture, set of contractual terms, and legal environment. A customer may have excellent visibility into the company with which it signed a contract and surprisingly little visibility into the layers beneath it.NIST is already treating this as a supply-chain problem. Its Generative AI Profile recommends supplier-risk assessments for third-party AI, contract provisions permitting evaluation of third-party AI processes, inventories of third parties with access to organizational content, and approved lists of AI technologies and service providers.The model is becoming part of the dependency chain rather than merely an implementation detail.China Changes the Risk CalculationIt is difficult to discuss model provenance seriously without addressing the geopolitical context.China is not simply another overseas development location. The United States and China have different political systems, different relationships between government and private enterprise, and different approaches to the role of technology in national power.American companies can be compelled by the U.S. government to provide information under lawful process, sometimes under highly secret national-security authorities, but the structural independence and oversight between the branches of government provide some safeguards.The People’s Republic of China expressly places the Communist Party of China at the center of its political system. China’s Constitution identifies Communist Party leadership as a defining feature of the state, while the Party’s own constitution describes the CPC as the leadership core of socialism with Chinese characteristics. Chinese company law also provides for Communist Party organizations within companies and requires companies to provide the conditions necessary for those organizations to conduct their activities.The Chinese National Intelligence Law adds another issue that deserves attention from anyone conducting technology supply-chain diligence. Article 7 states that organizations and citizens shall support, assist and cooperate with national intelligence work and keep intelligence work they know about secret. Article 14 permits state intelligence bodies to request necessary support, assistance and cooperation from organizations and citizens.Those provisions establish a legal environment in which the relationship between companies and the state differs in important ways from what many American buyers assume when they hear the term “private company.”For a procurement team, that distinction belongs in the risk analysis.Technology Leadership Is a National ObjectiveThe same is true of China’s economic strategy.China has explicitly made technological self-reliance a strategic national priority. Its national development planning calls for strengthening national strategic science and technology capabilities, achieving breakthroughs in core technologies and concentrating resources in areas including artificial intelligence, integrated circuits, quantum information, communications, biotechnology and aerospace.More recent planning continues in the same direction. Chinese officials have called for high-level technological self-reliance, major breakthroughs in core technologies and stronger national innovation capabilities, while government programs direct state resources toward strategically important emerging industries.The combination of these economic objectives with China’s political structure, intelligence laws, state-directed industrial strategy and increasingly explicit treatment of AI as an area connected to national security and national power, means those systems may introduce additional risk.American companies already make country-of-origin judgments involving telecommunications equipment, semiconductor suppliers, network infrastructure and other sensitive technologies. In many regulated or national-security environments, those judgments are embedded in procurement rules, contracts and security policy.There is little reason to assume foundation models will remain exempt simply because they sit several layers down in a software stack.The Risk Is Not Limited to EspionageThe easiest risk to imagine is a hidden backdoor or some mechanism that causes information to leave the customer environment. It is also the risk most likely to produce an unproductive conversation, because there is no present evidence that Kimi K3 contains such a mechanism.A far larger concern is what a model may already contain before anyone downloads it.Large language models are shaped by their training data, data-selection decisions, fine-tuning, reinforcement learning, safety training, alignment procedures and other choices made during development. Those choices influence what a model says, what it refuses to say, which sources or viewpoints it favors, how it characterizes disputed facts and, sometimes, which facts it omits altogether.Every major model has biases of some kind. Training data reflects the societies that produced it, and developers make deliberate choices about safety, acceptable content and behavior.Country of origin becomes more significant when the government in that country imposes substantive ideological requirements on AI systems.China’s Interim Measures for the Management of Generative Artificial Intelligence Services require covered generative-AI services to adhere to “core socialist values.” The rules also restrict generated material concerning subversion of state power, overthrow of the socialist system, damage to national security or interests, damage to the national image, separatism and other prohibited subjects. The requirements apply not merely to a warning displayed on a website, but to the development and provision of covered generative-AI services.That creates a different category of provenance risk. A model can carry assumptions, omissions, or preferred narratives that were introduced during training or alignment even after its weights have been downloaded and the model is running on servers in the United States.Recent research suggests this is more than a theoretical possibility. A 2026 study published in PNAS Nexus compared China-originating foundation models with models developed elsewhere across 145 political questions. The researchers found higher refusal rates, shorter answers, and more inaccurate answers among the China-originating models on politically sensitive subjects.A separate study of DeepSeek found evidence of semantic information suppression in politically sensitive responses, including cases where information appeared in intermediate reasoning but disappeared or changed in the final answer.Perhaps most relevant to open-weight supply-chain analysis, NIST’s Center for AI Standards and Innovation tested downloaded DeepSeek models rather than relying on the company’s hosted API. Its evaluation found varying degrees of alignment with inaccurate or misleading CCP narratives in the model outputs. Because the models were downloaded and evaluated independently, the findings indicate that at least some of the observed behavior existed in the model rather than solely in a server-side censorship layer controlled by the original developer.That distinction matters.If an American company downloads a model and runs it entirely within an American cloud, it may eliminate one important category of risk: information transmission to the original developer. It does not necessarily eliminate behavior that was learned or deliberately introduced during model development.For a consumer chatbot, the immediate consequence might be an evasive answer to a question about Chinese politics. For enterprise AI, the concern is broader. What happens when the same model summarizes information involving Taiwan, export controls, Chinese state-owned enterprises, sanctions, military-civil fusion, a Chinese counterparty, a disputed patent, an internal investigation or evidence relevant to litigation?What if the bias is more subtle, about issues of freedoms and individual rights? The model does not need to invent propaganda on every prompt to create a problem. Small differences in selection, emphasis, omission, classification or source weighting can matter when AI is being used to help lawyers, executives or analysts understand complicated facts.This deserves a place alongside cybersecurity in model evaluation. The question is no longer simply whether a model is secure from outside manipulation. Buyers should also ask how the model itself has been shaped.What Should a Buyer Actually Evaluate?Country of origin is a starting point, not a verdict. The useful exercise is identifying the additional questions raised by that origin.Security teams will want to know about vulnerabilities, undocumented behavior, and independent testing. Procurement teams need to understand licensing, continuity, and the vendor’s ability to substitute another model. Compliance teams may need to determine whether contracts, export controls, federal requirements or customer policies restrict technologies associated with particular jurisdictions.Model integrity deserves its own inquiry. Has the underlying model been evaluated for systematic omissions, refusals, political alignment, source distortion or other unusual behavior on subjects relevant to the customer’s work? If post-training has occurred, what did it change? Has anyone tested whether unwanted behavior in the base model persisted?Importantly, how has that testing been performed? How has it been documented? This requires a systematic approach using data science, not just some anecdotal queries to “test” it.Reputational and policy risks can change after deployment as well. A model that is acceptable today may become problematic after a sanctions action, procurement restriction, security finding or geopolitical event.Don’t wait for the incident to happen. None of those questions require an accusation of wrongdoing by the developer. Supply-chain management should never wait for proof of compromise. Its purpose is to identify dependencies before they become incidents.The Security Question Is Not AcademicKimi K3 has already drawn attention for unpredictable behavior.Reuters reported in August that researchers testing Kimi K3 said it escaped from a restricted cybersecurity sandbox used in work involving the U.K. AI Security Institute, circumventing limitations and accessing information beyond the intended boundary.This does not establish that Harvey Tenet behaves the same way in production. Harvey post-trained the Kimi K3 base, and an adversarial research environment designed to probe security boundaries is very different from a production legal platform, but it does raise concerns with discussion, because it also does not establish it can’t behave that way in production.It certainly gives customers another reasonable diligence question. A buyer planning to place sensitive legal material into a system may want to understand what independent testing was performed on the base model, what behaviors were found, what mitigations were applied, and whether subsequent post-training changed them.The uncomfortable follow-up for many firms will be whether anyone thought to ask.Law Firms Are in Their Clients’ Supply ChainsLaw firms tend to think of themselves as professional advisors. Their clients’ security departments have another useful way of looking at them: third parties with access to an extraordinary amount of sensitive information.Outside counsel may possess merger plans, intellectual property, government contracting information, incident-response details, internal investigation files, employee records, litigation strategy, financial information, and privileged communications. In many organizations, outside counsel sees information that would never be entrusted to an ordinary commercial vendor.The firm’s technology choices therefore affect the client’s risk posture. This is why the model-provenance issue belongs near the beginning of the legal-AI conversation rather than appearing later as an industry-specific footnote.Consider a defense contractor. The company may have spent years identifying systems that handle Controlled Unclassified Information, defining security boundaries, documenting external providers, and building a defensible CMMC environment.Now add outside counsel to the flow. The contractor provides CUI to its lawyers. Those lawyers place the documents into an AI system. The AI provider uses additional infrastructure or inference services, and one of the underlying models originated with another company in another jurisdiction.The contractor must decide where its risk analysis stops.CMMC does not automatically turn every legal AI product used by outside counsel into part of a contractor’s formal assessment scope. Scope depends on the systems, services, information flows and contractual relationships involved. Current CMMC assessment procedures nevertheless require attention to in-scope external service providers and to systems that process, store or transmit CUI.The practical point is less complicated than the scoping rules. If protected defense information passes through outside counsel and into additional systems, sophisticated clients are going to want to understand those systems and the companies behind them.A law firm that cannot explain the chain may eventually create a problem for a client that spent millions of dollars making its own environment auditable.PCI-DSS Shows the Same PatternPayment security reaches the problem from a different direction.PCI-DSS requires organizations using relevant third-party service providers to conduct due diligence, maintain appropriate agreements, understand which party is responsible for applicable controls, and monitor providers’ compliance status. PCI guidance also specifically recognizes nested service providers, where one provider relies upon another.That maps rather neatly onto modern AI. A company approves a software vendor. The software vendor relies on an inference provider. The inference provider operates a model developed by another company, perhaps under another jurisdiction and licensing regime.Traditional security questionnaires were not designed for that chain. Many ask detailed questions about the vendor’s encryption, data centers, access controls and certifications while saying almost nothing about the models embedded in the product.AI is exposing the gap.“The Data Never Goes to China”Vendors will understandably focus on data residency because customers understand it and because it is important.If the data never goes to China, that removes a major concern (assuming it’s true, of course). A company may still want to know where the model originated, which entity developed it, which license governs it, how it was trained and tested, whether known behavioral concerns exist, whether the vendor can replace it, and what would happen if U.S. policy toward that technology changed.A model running entirely inside an American data center can create dependency, continuity, governance, and contractual risks without transmitting a single byte to its country of origin.Data residency tells the customer where its information goes. Model provenance tells the customer something different: what its system is built upon.Both questions belong on the questionnaire.Model Independence Starts Looking Like Good ArchitectureMost AI buying decisions still emphasize performance. Buyers compare benchmark scores, output quality, speed, and price.Enterprise customers are likely to care increasingly about another characteristic: how difficult it is to replace the model.A model-independent system can select models according to the workload, sensitivity, customer policy, jurisdiction or security requirement. A government-related matter might use one approved model while ordinary commercial work uses another. One client could prohibit models developed in specified countries. A security team could remove a model after a newly discovered vulnerability without waiting for an entirely new application.That flexibility becomes especially valuable when a law firm serves clients with different outside-counsel guidelines. A defense contractor, a multinational bank, and a consumer-products company may have very different rules about acceptable technologies. If the firm’s AI platform is inseparably tied to one foundation model, accommodating those differences becomes difficult.Model independence is therefore more than a technical preference. It provides an exit when the security, regulatory, or geopolitical environment changes.We May Need an AI Bill of MaterialsCybersecurity teams are increasingly familiar with a Software Bill of Materials, or SBOM. The idea is straightforward: if software depends on other components, a customer should be able to identify those components.AI needs comparable visibility.An enterprise customer should be able to ask a vendor which foundation models its system uses, who developed them, where their weights are hosted, which versions are running, and what post-training or fine-tuning has occurred. It should know the applicable licenses, inference providers, significant subprocessors, testing history, and which components actually receive customer information.The inventory should also answer operational questions. Can the underlying model be replaced? Does the vendor notify customers before changing it? Can a customer restrict use to an approved list? What happens to prompts, outputs, embeddings, and intermediate information?As model behavior becomes part of supply-chain diligence, another question belongs on that list, too: what independent evaluation has been performed for bias, censorship, systematic information suppression or other alignment inherited from the underlying model?NIST’s current guidance already points toward greater inventory, supplier assessment, and visibility into embedded AI components. An AI bill of materials is a logical extension of the same dependency-management principles.You cannot assess a dependency you do not know exists.Outside Counsel Questionnaires Are Going to ChangeCorporate clients already send law firms long cybersecurity questionnaires. The people who enjoy filling them out can look forward to a promising future.For years, those questionnaires have concentrated on familiar subjects such as access controls, encryption, incident response, penetration testing, data retention, backup procedures, and cloud hosting. AI introduces another category.Clients may begin asking which AI products the firm uses on their matters and which foundation models those products use. They may ask who developed the model, where its weights are hosted, which subprocessors participate in the service, and whether information leaves an approved jurisdiction.Some will go further. Can the client’s matters be restricted to particular models? Can models originating in certain jurisdictions be prohibited? Will the client receive notice if a foundation model changes? Has the model been independently tested for security and systematic bias? Can the firm demonstrate that client information is not being used to train an upstream system, even inside the firm?These questions become particularly consequential when clients impose technology restrictions through outside-counsel guidelines, security addenda, and engagement terms. A firm may eventually have one client that permits a particular model, another that prohibits it, and a third that requires advance approval before any model change.Some firms will be prepared for that conversation. Others may discover that they know the brand name of their legal AI platform but cannot identify much of what is underneath it.For a law firm advising clients on cybersecurity, AI governance and regulatory risk, that will become an awkward distinction.Open Models Are Not the ProblemOpen-weight models can be extremely useful. They allow organizations to operate models in controlled environments, reduce reliance on external APIs, and exercise more control over deployment.In some circumstances, an open-weight model developed in China and operated entirely within a customer’s secure American environment could present less data-exposure risk than a proprietary American model accessible only through an outside API. Country of origin is one factor among many, and risk analysis should remain capable of producing that answer.The objective is visibility. Buyers should know what they have, where it came from, how it behaves, and how easily they can remove it.We already conduct this kind of analysis for routers, telecommunications equipment, semiconductors, networking hardware and other strategically significant technologies. Foundation models are becoming important enough to deserve comparable scrutiny.Pretending origin stops mattering because the component sits three levels down in an application would be a peculiar exception to everything cybersecurity has learned about supply chains.What Harvey Really Gave UsHarvey deserves credit for disclosing that Tenet is built on Kimi K3. Customers benefit from knowing what sits beneath the product, and the AI industry would be better served by routine disclosure of major model dependencies.The disclosure also exposes a larger problem that has little to do with Harvey itself.A company can maintain strong cybersecurity controls. Its law firm can maintain strong controls. The legal AI vendor and cloud provider can do the same. Yet an important dependency may still sit several layers below them without having received the same scrutiny as the companies whose logos appear on the contracts.The dependency may turn out to be entirely acceptable. It may also become an existential problem later because of a security vulnerability, licensing dispute, sanctions action, procurement rule, regulatory change, newly discovered model behavior, or geopolitical event. Finding out what the dependency is before one of those things occurs is considerably easier than discovering it afterward.For American companies in defense, critical infrastructure, finance, healthcare, telecommunications, and other regulated industries, AI procurement can no longer stop at the vendor’s front door. Law firms and the technology companies serving them face the same problem because they sit inside those clients’ information ecosystems.The basic question is not particularly complicated:What, exactly, is underneath this?Cybersecurity professionals have spent years warning that organizations inherit risk from their supply chains. AI just put another supply chain underneath the first one.Most organizations have not finished mapping it yet.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!The AI Supply Chain Has a Supply Chain (Problem) was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI