There's a pattern that repeats throughout history with striking precision. Someone arrives with superior tools of documentation, records what entire communities refined over generations, and that knowledge later reappears in another context, under another name, generating wealth for whoever captured it. What distinguishes the past from the present isn't the logic, but the vocabulary wrapped around it.
The relevant question isn't whether indigenous knowledge is being integrated into AI models. That's already happening. What matters is under what conditions it happens, with what forms of consent, and who retains the benefits when that knowledge becomes a competitive advantage for a language model. Data encodes power.
The most direct historical parallel comes from colonial botanical expeditions. The Royal Botanical Garden in Madrid, Kew Gardens in London, and Humboldt's expedition systematized knowledge that indigenous communities had perfected over generations: medicinal plants, cultivation techniques, the properties of natural compounds. That knowledge fed pharmaceutical industries whose profits never reached Amazonian or Mesoamerican communities. The 1992 Convention on Biological Diversity sought to correct this; the 2010 Nagoya Protocol established frameworks for access and benefit-sharing. More than a decade later, progress remains limited. The gap between formal recognition and actual compensation persists.
The structure is always the same. Knowledge generated collectively over generations gets systematized by outside actors with greater resources, turned into an economic asset or intellectual property, and the originating communities are left out of the benefit cycle. Decades later comes symbolic recognition, without concrete redistribution. What's happening with AI replicates that same flow, just at a multiplied speed of extraction.
Anyone who has worked with data systems knows that digitizing traditional knowledge carries implications that aren't immediately obvious. When indigenous botanical knowledge, forest management practices, or ancestral climate-prediction systems become text used to train a model, the community context gets diluted in the encoding. A model doesn't learn the Kayapó people's knowledge of fire management. It absorbs statistical regularities between concepts. Collective authorship, oral transmission, and the ritual context that gives that information meaning disappear in tokenization.
This is where a technical problem emerges that rarely surfaces in governance debates. There are still no mechanisms applied at scale to trace the provenance of indigenous knowledge within large models. The CARE principles — centered on Collective benefit, Authority to control, Responsibility, and Ethics — emerged as a counterweight to the FAIR principles, which prioritize accessibility and reuse. CARE, however, remains an academic proposal. There's no mandatory labeling to flag which segment of a model was trained on Mapuche knowledge about medicinal plants and thus requires compensation to that specific community. That infrastructure is missing because the incentives to build it simply aren't aligned.
International institutions position themselves as guardians of indigenous knowledge while the frameworks that actually matter — data ownership and economic compensation — remain without binding force. UNESCO, the CBD, and various UN bodies produce declarations, recommendations, and best-practice documents. Integrating indigenous knowledge into national climate adaptation plans is a suggestion, not a mandatory requirement in most jurisdictions. Researchers have spent years pointing out that this symbolic governance can function as cover for real extraction: while the conversation stays focused on recognition and inclusion, it avoids the question of who controls the models trained on that knowledge.
Technically, it's feasible to design alternatives, even if politically inconvenient. Informed consent can be encoded into contracts that establish conditions of use, limits, and automatic compensation. Licenses for collective data can include specific clauses, similar to those that differentiate uses under Creative Commons. Federated architectures would let communities retain local control while still contributing to broader models. These approaches face real limitations, especially where digital divides run deep. There are no clean solutions. Still, the difference between technically difficult and technically impossible matters, because the dominant narrative tends to erase that distinction to justify inaction.
The post-COP30 landscape adds both urgency and irony. Commitments on fossil fuels fell short of what communities demanded, while at the same time, the knowledge those same communities hold about adaptation and ecosystem management is exactly what models need to improve their environmental predictions. Those who know the most about living in balance with ecosystems that the rest of the world is damaging are those with the least control over how that knowledge becomes a tech product. That contradiction rarely shows up in summit communiqués.
What's telling about this moment is how quickly the language of inclusion can hollow out concrete demands. Participatory governance, valuing traditional knowledge, co-creation with communities: these phrases can mean symbolic consultations that end in the same extractive outcome, or they can mean real control over data, veto power over specific applications, and a share in economic benefits. The difference isn't in the rhetoric — it's in the technical and legal details that few people in power actually want to negotiate.
The colonial pattern of scientific appropriation didn't need conscious villains. It only needed institutions with incentives to extract, and no counterweight to stop them. Today, companies incorporating traditional knowledge into their models don't need malicious intent either. They just need the frameworks to stay recommendatory, the communities to lack legal recourse to dispute ownership, and public debate to remain at the level of principles instead of descending into actual contracts. That's exactly what's happening.
I don't have a clear answer for how to fix this in the short term, and it would be dishonest to present a tidy roadmap. What does seem evident is that the convergence between AI governance and indigenous knowledge will generate intellectual property tensions that current institutions aren't equipped to resolve. The pressure will have to come from the communities themselves developing data sovereignty frameworks, and from those willing to name extraction for what it is, even when it comes wrapped in the language of cultural preservation.
How do we make sure that, this time, ancestral knowledge benefits first those who have safeguarded it for generations?
Sources:
1. Convention on Biological Diversity — Nagoya Protocol on Access and Benefit-Sharing (2010), CBD Secretariat
2. Carroll, S.R. et al. — "The CARE Principles for Indigenous Data Governance," Data Science Journal (2020)
3. Wilkinson, M.D. et al. — "The FAIR Guiding Principles for scientific data management and stewardship," Scientific Data (2016)
4. Mignolo, W. — Local Histories/Global Designs: Coloniality, Subaltern Knowledges, and Border Thinking, Princeton University Press (2000)
5. Kukutai, T. & Taylor, J. (eds.) — Indigenous Data Sovereignty: Toward an Agenda, ANU Press (2016)