Denmark takes the lead in the race for small language models
Just a few years ago, warnings sounded that Europe risked falling behind in the global AI race. Now, a research team from the University of Southern Denmark (SDU) and Ordbogen A/S, the company behind the AI infrastructure Odincore.ai, has created the language model Mimir – named after the guardian of the well of wisdom in Norse mythology.
When researchers working on the Danish Foundation Models (DFM) project first began developing Danish language models in earnest, the outlook was not exactly optimistic.
The largest and most advanced models were being developed by American and Chinese tech giants, which had – and still have – access to enormous amounts of data, computing power and billion-dollar investments. Denmark had neither the same volumes of data nor the same computing resources.
The risk was that Denmark – and Europe – would have to settle for becoming customers of companies on the other side of the Atlantic or in China, with all that this entails.
But over the summer, the researchers in the DFM project developed a new language model with one billion parameters that can now compete with those of the tech giants.
As Professor Peter Schneider-Kamp from the University of Southern Denmark puts it:
– For the first time, we have developed a model that is not only relatively good at Danish, but which, given its size, is currently the best model in the world.
A Danish lightweight takes on the tech giants
One billion parameters may not immediately sound like a small model. But in the world of modern AI, it belongs firmly in the lightweight class.
What is interesting, therefore, is not simply how well the model performs. It is how well it performs relative to its size.
The researchers have tested the model using standardized evaluations in which language models are assessed on, among other things, their ability to handle knowledge, mathematics, programming, translation and various language tasks.
According to Schneider-Kamp, the Danish model performs better than leading small models from, among others, Google's latest Gemma models and Alibaba's Qwen model family.
– In English, we have to go up to a model that is around four times larger to reach the same level. And in Danish, even models up to nine times its size cannot keep up, he says.
The model also performs well in mathematics and programming, where it can compete with models four to five times larger.
This does not mean that Denmark has suddenly built a competitor to the largest versions of ChatGPT. Those models belong to an entirely different weight class.
– ChatGPT is around 1,000 times larger. But we can compete with the best models they can build at this size. In fact, we can do better than they do, says Schneider-Kamp.
So, it is not the heavyweight contest that the Danish researchers have won. But in the lightweight class, they are currently leading the race.
A lead that may not last long
It is worth paying attention to the words right now.
Artificial intelligence is developing so rapidly that records can have very short lifespans.
– That is also why we need to get it out there quickly, because we simply do not know what is going to happen. Maybe in a week or in two months, someone will release a model that is even better, says Schneider-Kamp.
Even so, the result changes something fundamental about the narrative surrounding Europe's position in the AI race.
A few years ago, the question was whether countries such as Denmark had any chance at all of keeping up with the enormous American and Chinese investments.
The DFM project has now demonstrated that it is possible to compete at world-class level without copying the enormous scale of the tech giants.
The secret is not more – but, quite literally, less
An important part of the explanation lies in the way the model has been trained.
The classic recipe for modern language models has, roughly speaking, been: more parameters, more computing power and enormous amounts of text.
But AI research is also increasingly moving towards finding ways to get more out of less. This is the development that the Danish model takes advantage of. According to Schneider-Kamp, the model performs better than models trained on up to 100 times as much data.
This is particularly interesting for a relatively small language such as Danish. English-language models can be trained on almost unimaginable quantities of text from the internet. A language spoken by around six million people does not have that luxury.
– We do not have that much data, and certainly not that much Danish data. But because we do not need as much data, suddenly we can compete after all, says Schneider-Kamp.
According to him, around 22 per cent of the model's training data is Danish. A further five per cent concerns translation to and from Danish.
The model is also presented with specific tasks: summarize this text/find the errors/continue the story/translate between Danish and English/solve mathematical problems, and so on.
This more targeted training is one of the reasons it is possible to get so much out of a relatively small model – one that is, in fact, small enough to live on a laptop.
And its small size has another significant consequence. The model requires far less computing power to use than the enormous models that many people associate with generative AI. First, this means lower energy consumption. Second, it means that the data a user enters into the model does not need to leave the computer.
– Then the only electricity you use is what your own machine consumes. And no data is sent anywhere else, explains Schneider-Kamp.
This potentially opens the door to applications where organizations do not want to send data to external AI services, and where the cost or energy consumption of large models would otherwise be a barrier.
AI does not have to consume all our electricity
The question of size also touches on one of the major concerns that has accompanied the AI boom.
When ChatGPT broke through, it was followed by forecasts of exploding demand for data centres, chips and electricity. Fortunately, however, developments have also moved in another direction – models have become more efficient.
The new Danish model is an example of this development.
– It is actually quite remarkable and reassuring. It gives us some hope, says Schneider-Kamp.
His point is not that the issue of AI's energy consumption has thereby been solved. After all, we are using the technology more and more. But if new models can solve the same tasks at a fraction of the size and using far less training data, it changes the equation. And the researchers are unlikely to have found the most efficient method yet.
From AI dependency to digital sovereignty
For Schneider-Kamp, however, the most important thing is not a benchmark or a position on a ranking.
It is primarily about whom Denmark and Europe depend on.
Today, some of the world's most advanced AI systems are provided by a very small number of American and Chinese companies. This means that access, prices, technical limitations and, ultimately, the rules under which the models operate are largely determined outside Denmark.
A model that can be downloaded and run locally changes that relationship.
– It means that we can have data sovereignty. That we can be independent of Americans and Chinese companies that might switch off their models if they decide Europeans should no longer have access to them, says Schneider-Kamp.
He also points to the economics. The largest AI models are extremely expensive to develop and operate. Small, efficient models can therefore help make the technology accessible to more people and reduce dependence on the subscription services controlled by the tech giants.
– It is better for the environment, while at the same time helping to democratise AI in a way, he says.
Denmark is leading the race – now the challenge is to keep up the pace
The Danish Foundation Models project was launched partly out of concern about what happens if the Danish language, Danish values and Danish institutions become little more than a footnote in an AI development dominated by the world's largest countries and companies.
That concern has not disappeared, but the development of the new model shows that the conclusion does not necessarily have to be that a small country such as Denmark simply cannot keep up.
Perhaps we do not always need the largest models, the largest volumes of data and the largest data centres. Sometimes technological competitiveness is also about finding a smarter way of doing things.
And right now, the situation has almost been turned on its head. Denmark is no longer merely trying to catch up with the others. In the lightweight class – at least until the next model is released – it is the others who will have to try to catch up with Denmark.
Contributors to the development of Mimir: Peter Schneider-Kamp, Jacob Nielsen, Lukas Galke Poech, Gianluca Barmina, Mogens Henrik From, Andrea Blasi Núñez, Kenneth Enevoldsen, Annemette Brok Pirchert, Stine Lyngsø Beltoft, Torben Blach, Sofie Helene Bruun, Oliver Kinch, Rasmus Larsen, Dan Saattrup Smart and Kristoffer Laigaard Nielbo. Click on the image to view a larger version.
FACT BOX: Danish Foundation Models
Danish Foundation Models (DFM) is a Danish research collaboration working to develop open language models with a particular focus on the Danish language, culture and society, funded by the Ministry of Science, Higher Education and Digital Affairs
The project brings together researchers from, among others, the University of Southern Denmark, Aarhus University, University of Copenhagen and the Alexandra Institute, as well as ordbogen.com, which has also played a key role in the project.
The work involves not only developing the language models themselves, but also Danish training datasets and benchmarks used to test how well the models handle the Danish language and Danish contexts. A key objective is to build Danish expertise and develop AI models that can be examined, adapted and run within frameworks that provide greater control over technology and data.
DFM has already released several models and datasets. The new model with one billion parameters is the project's strongest result to date. The model is available on Hugging Face, an open-source platform for AI developers.
Meet the researcher
Peter Schneider-Kamp is a professor at the Department of Mathematics and Computer Science at the University of Southern Denmark in Odense.
