Wikipedia is getting a machine-friendly makeover. Wikimedia Deutschland has unveiled the Wikidata Embedding Project, a new database designed to make the site’s vast knowledge base easier for AI models to use.
The system applies vector-based semantic search, a way for computers to grasp the meaning and relationships between words, across nearly 120 million entries from Wikipedia and sister projects.
It also supports the Model Context Protocol, a standard that helps AI systems talk directly to data sources.
Together, these upgrades allow large language models to query Wikipedia with natural language, making it far more useful for retrieval-augmented generation (RAG) systems.
Built with Jina.AI and IBM-owned DataStax, the database goes beyond keyword searches.
A query for “scientist” might return names, translations, images, and related concepts such as “researcher” or “scholar.”
Project manager Philippe Saadé stressed its open ethos: “Powerful AI doesn’t have to be controlled by a handful of companies.”