More Resource-Efficient AI

Vicky Kalogeiton, a researcher at the École Polytechnique’s Computer Science Laboratory (LIX*), has just received a Starting Grant from the European Research Council (ERC). The goal of her research project FLASH is to make generative image AI more efficient, more accessible, and more sustainable.
03 Sep. 2026
Research, Awards, IA et Science des données, LIX, École polytechnique

“The more data, the larger the model, the better the results.” This, in broad strokes, is the mantra that has guided the evolution of artificial intelligence models for several years. The emergence and subsequent improvement of AI systems such as ChatGPT, Claude, and Mistral, trained on extremely large datasets using substantial computational resources, illustrate this paradigm. But it leaves one scientific question unanswered: Is it possible to build models that deliver high-quality, controllable results by taking a more efficient approach?

From Model Size to Efficiency

In a world with finite resources, where sustainability is a key challenge for the future, this issue is of great importance. This is the motivation behind Vicky Kalogeiton’s research project, funded by the European Research Council. Its name: FLASH (From scaling to efficiency laws for visual synthesis). The goal is to shift from the paradigm of scale to that of efficiency for visual generation (images, videos, etc.).

 “There are three areas that need to be optimized: the data used to train the model, the model itself that performs the calculations, and finally the interactions, that is, the back-and-forth between the user and the AI to achieve the desired result,” explains the LIX researcher.

In preliminary work, Vicky Kalogeiton and her colleagues demonstrated that it was possible to use 1,000 times less training data while still achieving very good performance (1). How can we reduce this size while maintaining the same performance? To what extent do we need diverse images? What is the impact of having duplicate images or of the type of image (natural landscapes versus abstract figures)?

Data-Driven Training

The next step is to optimize the model by “guiding” it with data. If training images are typically annotated (for example, “a snow-covered mountain range”), current training approaches do not sufficiently take into account the structure and characteristics of the data itself, which can lead to inefficient use of resources. “The goal will be to extract additional information from the data in order to better guide the model’s training.,” emphasizes Vicky Kalogeiton. 

For a user, the final result is often obtained after multiple “prompts,” or interactions, with the AI. How can we make these exchanges shorter while improving accessibility? The AI could ask for clarification when necessary before generating an image, or even provide a rough sketch first.

Applications in Robotics

In addition to image and video generation, FLASH will also explore embodied AI applications in robotics, where data, computation, and interactions are particularly constrained. For example, using images of a robotic arm’s surroundings, an AI model must be able to understand its environment and provide the sequence of actions necessary for the arm to grasp an object.

Finally, FLASH paves the way for more efficient embedded models that can run directly on mobile devices, reducing reliance on external servers while improving privacy and lowering energy consumption.

 

(1) Lucas Degeorge, Arijit Ghosh, Nicolas Dufour, David Picard, and Vicky Kalogeiton. How far can we go with imagenet for text-to-image generation ? In arXiv preprint arXiv:2502.21318 , 2025

 

*LIX: a joint research unit CNRS, École Polytechnique, Institut Polytechnique de Paris, 91120 Palaiseau, France

Back