AI · Web3 · Tech trends and insights at a glance
AI · Web3 · Tech trends and insights at a glance
When Meta, DeepSeek, and Google invoke 'open source AI,' what they actually release is model weights — the outputs of training processes whose data, pipelines, and reward models remain proprietary. This gap between open weights and open source is not semantic nitpicking: it determines who can audit AI systems for bias, who qualifies for regulatory exemptions under the EU AI Act, and whether 'openness' is a transparency principle or a market strategy.
The phrase 'open source AI' has become one of the most productive pieces of marketing language in recent technology history. Meta deploys it each time a new Llama model ships. DeepSeek borrowed its credibility to position R1 as a challenge to Western AI dominance. Google frames Gemma as an act of openness; Mistral presents itself as Europe's answer to a closed-AI future. The proclamation that open source is 'the path forward' has become competitive ritual, repeated so frequently that scrutiny of what it actually means tends to dissolve in the noise.
What these companies are releasing, in practice, is model weights: the numerical parameters that encode a trained neural network's learned behavior. This is not nothing. Access to weights allows researchers and developers to run inference, fine-tune models for specific tasks, and build applications without dependency on a proprietary API. But it is also considerably less than what 'open source' has meant in software for the past three decades, and the gap between the two carries consequences that now extend from academic reproducibility into the architecture of AI regulation itself.
The Open Source Initiative, the body that has defined open source software since 1998, published its Open Source AI Definition in late 2024. The standard is explicit: an AI system qualifies as open source only if the training data, data processing code, and training scripts are made available alongside the model weights, under terms permitting unrestricted use, modification, and redistribution. On this definition, none of the models most commonly described as 'open source' actually qualify.
Meta's Llama 4 license imposes commercial use restrictions beyond a certain deployment threshold and limits how derivative models can be redistributed — terms that would fail OSI's criteria even if training data were available, which it is not. DeepSeek's R1 release withholds the composition of training corpora and imposes comparable commercial constraints. Google's Gemma releases technical reports that sketch the shape of the training process without disclosing dataset provenance or the reward model design used in reinforcement learning from human feedback. The more accurate term for all of these is 'open weights' — a genuinely useful category, but a categorically different one from open source.
The distinction is not pedantry. A trained language model is not analogous to compiled software whose source code has been lost. The model's behavior is constitutively shaped by the data it was trained on, the fine-tuning objectives, the safety filters applied, and the human preference data embedded through RLHF. When that context is withheld, external researchers cannot reproduce the model's behavior, audit its biases, trace safety failures to their origin, or independently verify vendor claims. Publishing the pattern without disclosing the provenance is a form of opacity dressed in the language of transparency.
The stakes of this definitional contest are no longer confined to academic argument. The EU AI Act establishes tiered obligations for AI systems based on risk classification, and 'open source' status functions as a partial exemption criterion for certain transparency and documentation requirements. If weight-sharing alone were sufficient to claim that exemption, developers could sidestep data documentation mandates, risk management frameworks, and post-market monitoring obligations by releasing a model file and invoking the open source label. The European Commission's AI Office has signaled awareness of this loophole and is working toward implementing guidance that interprets 'open source' for AI more narrowly than simple weight availability — but legal clarity has not arrived, and the ambiguity currently benefits incumbents.
Competition law presents a parallel dynamic. Meta's decision to release Llama weights at no cost looks like generosity from one angle and like a market-foreclosure strategy from another. By making a powerful base model freely available, Meta sets the gravitational center of the fine-tuning ecosystem around its own architecture choices, training defaults, and acceptable-use policies — while retaining exclusive control over the training infrastructure and data pipelines that would be required to replicate the model from scratch. The ecosystem builds on Meta's foundation without being able to inspect, challenge, or reproduce it. The UK Competition and Markets Authority and the European Commission's DG COMP have both begun examining whether foundation model strategies employed by large technology companies raise concerns under existing competition frameworks, with the question of what 'openness' actually entails sitting at the center of that inquiry.
There are models that take the more demanding interpretation of openness seriously. The Allen Institute for AI's OLMo series releases training data in the form of the Dolma dataset, preprocessing code, training scripts, and intermediate checkpoints alongside weights — a standard of transparency that enables genuine reproducibility. EleutherAI's Pythia models were designed from the outset to support mechanistic interpretability research, with full reproducibility as an explicit design goal rather than an afterthought. These projects demonstrate that rigorous open source LLM development is technically feasible, even if it requires confronting the uncomfortable question of what data you actually trained on and accepting that others will scrutinize it.
The limiting factor is not capability but incentive structure. For companies whose model weights encode competitive advantage derived from proprietary data curation, internally-designed safety tuning, and training pipelines built at enormous computational cost, genuine openness eliminates the moat. Publishing weights while withholding everything else offers an approximation of open source's reputational benefit without its competitive cost. As long as regulators, journalists, and the broader public treat the two as equivalent, that calculus will hold.
The endgame of the definition war depends on which institutions have the authority and will to adjudicate it. If the EU AI Act's implementing rules enforce a rigorous definition that requires data transparency as a condition of open source exemption, the label will lose its value as a regulatory bypass. If antitrust investigators conclude that open weights paired with closed pipelines constitutes a form of ecosystem lock-in, the strategy will attract scrutiny rather than praise. At its core, the open source AI debate is about where power over these systems actually resides — in the hands of those who publish the numbers, or in the hands of those who understand, and exclusively control, how those numbers were made.
The Land-Permit Paradox of Korea's Chip Belt, When the Cluster's Boom Prices Out Its Own Engineers
Dongtan, Giheung, and Guri have been folded into Korea's land-transaction permit regime just as the AI chip capex boom reshapes the property market around the country's largest fabs. The very prosperity the cluster generates is raising the cost for the engineers it depends on to settle nearby. The real test of agglomeration may lie not in siting megafabs but in housing and labor mobility.
The Collapse of the Closed AI Moat and the Supply-Chain Paradox of Unverifiable Weights
DeepSeek-R1's open reasoning weights and Llamafile's single-file distribution are eroding the performance and distribution moats that closed labs once charged a premium for. Yet the same openness collides head-on with the gap exposed by the "250 samples to break an LLM" research: weight distribution that no recipient can verify. Democratized competition and accumulated security debt now sit on the same scale.
Forty-Year Yen Lows as the Hidden Subsidy Behind Japan's Chip Revival
As the yen slides into its weakest territory in four decades, Takaichinomics has entered uncharted monetary terrain. A cheap yen functions as a silent subsidy for Rapidus, Kioxia, and TSMC's Kumamoto fabs—yet the same currency inflates the cost of imported tools and materials and intensifies the talent war with Korea. The question is whether monetary policy can stand in for industrial policy, and what that means for Korea's memory champions.