Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

Jump to content

Kandinsky (artificial intelligence)

From Wikipedia, the free encyclopedia
Kandinsky
DeveloperSber AI
Release14 June 2022; 4 years ago (2022-06-14)
Stable release
Kandinsky 5.0 / 20 November 2025; 10 months ago (2025-11-20)
Operating systemWeb service
Type
LicenseProprietary (web service); open-source (model weights)
Websitekandinskylab.ai

Kandinsky is a series of generative artificial intelligence models built by Sber. The models turn text prompts into images and short videos.[1] Sber AI leads the work, with help from the Artificial Intelligence Research Institute (AIRI) and SberDevices. The painter Wassily Kandinsky gives the project its name.[2]

The first model shipped in June 2022. It used the autoregressive design Sber had built for ruDALL-E.[3] Five months on, Kandinsky 2.0 switched the family to multilingual latent diffusion across 101 languages.[4][2] Five months on, Kandinsky 2.0 switched the family to multilingual latent diffusion across 101 languages. A second big change came with Kandinsky 3.0 (November 2023), which dropped the intermediate "image prior" used in 2.x and fed Flan-UL2 features straight into the diffusion U-Net.[5] Video generation followed in late 2023, and the open-source Kandinsky 5.0 line in November 2025.[6][7]

Russian and Western tech press has called Kandinsky Sber's reply to DALL-E, Midjourney and Stable Diffusion.[8][2] The project has also pulled scrutiny from Russian politicians and from the country's IT regulator.[9][10]

History

[edit]

Sber's AI lab released ruDALL-E in November 2021. The 1.3-billion-parameter autoregressive model was a Russian-language port of OpenAI's DALL-E and it served the public through a web demo called rudalle.[3] Kandinsky 1.0 followed seven months later, on the same autoregressive bones but scaled to 12 billion parameters and trained on 179 million image–text pairs. The pipeline ran in three stages: candidate generation, ruCLIP-based reranking, then upscaling through Real-ESRGAN or a diffusion model.[3]

Sber unveiled Kandinsky 2.0 at its AI Journey conference on 23 November 2022. The new model dropped autoregression for latent diffusion, on the model of Stable Diffusion. Two multilingual encoders, XLM-R-CLIP and mT5-small, made room for prompts in 101 languages, and the 1.2-billion-parameter U-Net was trained on roughly a billion image–text pairs across 196 NVIDIA A100 GPUs over 14 days.[2][4] Kandinsky 2.1 in April 2023 cut the parameter count to 3.3 billion, swapped the VQGAN decoder for MoVQ and added an "image prior" module that mapped CLIP text embeddings into the image side before diffusion.[1]

Image resolution doubled to 1024×1024 with Kandinsky 2.2 in July 2023, alongside ControlNet-style editing. A short-clip video mode was added in October.[11] The next AI Journey, in November 2023, brought Kandinsky 3.0 - the first big architectural rebuild since 2.0. It removed the image-prior stage, swapped the text encoder for Flan-UL2 and replaced the U-Net interior with BigGAN-deep blocks.[12][5] A dedicated text-to-video model, Kandinsky Video, premiered at the same conference. A second EMNLP demonstrations paper, this time on Kandinsky 3, was accepted to the 2024 conference in Miami.[5] Distilled and updated variants - Kandinsky 3.1 in April 2024 (four-step inference, image-to-image) and Kandinsky 4.0 in December 2024 (twelve-second HD video) - followed without a full architectural change.[13][14]

Kandinsky 5.0 launched on 20 November 2025 with code and weights on GitHub. Three families were released together: Image Lite (6 billion parameters), Video Lite (2 billion) and Video Pro (19 billion, HD video with controllable camera motion).[6] The technical report puts the training corpus at roughly a billion images and 300 million videos, with reinforcement learning from human feedback used for post-training alignment.[7][15]

Release history

[edit]
Kandinsky releases
VersionDateArchitecture / notes
Kandinsky 1.0June 2022Autoregressive, 12 B parameters[3]
Kandinsky 2.0November 2022Latent diffusion, 1.2 B U-Net, 101 languages[4][2]
Kandinsky 2.1April 20233.3 B parameters, image prior, MoVQ decoder, 768×768; FID 8.03 on COCO-30K
Kandinsky 2.2July 20231024×1024, ControlNet-style editing; short video mode added October 2023
Kandinsky 3.0November 2023Flan-UL2 text conditioning, BigGAN-deep U-Net[12][5]
Kandinsky 3.1April 2024Distilled (4-step inference), image-to-image[13]
Kandinsky VideoNovember 20238 s clips, 30 fps, 512×512; keyframe + interpolation[16]
Kandinsky 4.0December 202412-second HD video[14]
Kandinsky 5.0November 2025Three families: Image Lite (6 B), Video Lite (2 B), Video Pro (19 B); open-source[6][15]

Controversies

[edit]

A Just Russia – For Truth leader Sergey Mironov took Kandinsky 2.1 to Prosecutor General Igor Krasnov in April 2023, asking the office to check the service against Russian law. Mironov said its outputs formed a "negative image" of Russia. He cited prompts about Russia and the Russian flag that came back without the national colours, a "Donbas is Russia" prompt that returned an image in Ukrainian colours, and a Z-patriotism prompt that returned what he called a "zombie-like creature". He accused Sber of reusing Western models without adapting them.[9]

That July, Vedomosti reported that Sber had written into the service's user agreement that it was not legally responsible for the images Kandinsky produced. The paper tied the disclaimer to Mironov's inquiry and to a wider Russian debate about liability for generative-model output.[17]

A Ministry of Digital Development expert council voted 20 - 4 against adding Kandinsky to the Unified Register of Domestic Software in February 2026. Rule 5 of the registry covers foreign control, source-code location and infrastructure independence. The deciding objection was that the user manual listed Telegram as the only sign-in channel, which the council member Andrey Vishnyakov called an "infrastructural dependence on an external service". Sberbank said it would revise the application and resubmit.[10]

Radio Free Europe/Radio Liberty reported in 2025 that some Sber AI staff working on Kandinsky left Russia in 2022 with an estimated 100,000 other IT specialists, after the invasion of Ukraine and Western sanctions cut off advanced chips and cloud services.[18] bne IntelliNews has likewise described Russian AI as advancing but still behind Western and Chinese competitors under sanctions.[19]

Censorship has surfaced in Sber's other generative products, not in Kandinsky itself. Meduza reported in January 2025 that Sber's chatbot GigaChat refused questions about the death of Alexei Navalny or the start of the war in Ukraine, returning a fixed line that "discussions on certain topics are temporarily restricted". The Chinese model DeepSeek, by comparison, answered the same questions.[20]

See also

[edit]

References

[edit]
  1. 1 2 Razzhigaev, Anton; Shakhmatov, Arseniy; Maltseva, Anastasia; et al. (December 2023). Kandinsky: An Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Singapore: Association for Computational Linguistics. pp. 286–295. doi:10.18653/v1/2023.emnlp-demo.25.
  2. 1 2 3 4 5 Yalalov, Damir (23 November 2022). "Sber AI has presented Kandinsky 2.0, the first text-to-image model for generating in more than 100 languages". Metaverse Post. Retrieved 10 June 2026.
  3. 1 2 3 4 "«Сбер» представил Kandinsky — ИИ-модель для генерации изображений по текстовому описанию на русском языке". 3DNews (in Russian). 14 June 2022. Retrieved 11 July 2023.
  4. 1 2 3 "Сбер показал нейросеть Kandinsky 2.0". RBC (in Russian). 23 November 2022. Retrieved 11 July 2023.
  5. 1 2 3 4 Arkhipkin, Vladimir; Vasilev, Viacheslav; et al. (November 2024). Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework. Proceedings of the 2024 EMNLP: System Demonstrations. Association for Computational Linguistics. pp. 475–485.
  6. 1 2 3 Arkhipkin, Vladimir; Korviakov, Vladimir; et al. (19 November 2025). "Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation". arXiv:2511.14993 [cs.CV].
  7. 1 2 "Sber Open Sources Russia's Most Advanced AI Models Under MIT Licence". Open Source For You. 24 November 2025. Retrieved 10 June 2026.
  8. ↑ "Next version of Kandinsky 3.0 by Sber". Channel 360 MEA. Retrieved 10 June 2026.
  9. 1 2 "Миронов попросил Генпрокуратуру проверить нейросеть «Сбера» Kandinsky из-за антироссийского контента". Kommersant (in Russian). 26 April 2023. Retrieved 10 June 2026.
  10. 1 2 "Власти не признали нейросеть Сбербанка российским ПО". CNews (in Russian). 6 March 2026. Retrieved 10 June 2026.
  11. ↑ "Сбер представил новую версию своей нейросети Kandinsky". Gazeta.ru (in Russian). 12 July 2023. Retrieved 13 July 2023.
  12. 1 2 "Сбер представил новую версию нейросети Kandinsky 3.0". TASS (in Russian). 22 November 2023. Retrieved 30 April 2024.
  13. 1 2 "Изобразительная нейросеть Kandinsky 3.1 стала доступна для всех пользователей". 3DNews (in Russian). 22 April 2024. Retrieved 30 April 2024.
  14. 1 2 "Сбер обновил нейросеть Kandinsky для создания видео". Hi-Tech Mail.ru (in Russian). 12 December 2024. Retrieved 14 December 2024.
  15. 1 2 "Sber unveils open-source AI suite for Russian language & media". ITBrief UK. 21 November 2025. Retrieved 12 June 2026.
  16. ↑ "В России появилась первая нейросеть, генерирующая видео". Telecom Daily (in Russian). 23 November 2023. Retrieved 27 September 2024.
  17. ↑ "«Сбер» и разработчики Kandinsky не будут отвечать за результат работы нейросети". Vedomosti (in Russian). 17 July 2023. Retrieved 12 June 2026.
  18. ↑ "As DeepSeek Rises, Russia Falls Behind On AI". Radio Free Europe/Radio Liberty. 7 February 2025. Retrieved 12 June 2026.
  19. ↑ "Russian AI accelerates but sanctions leave it trailing global leaders". bne IntelliNews. Retrieved 12 June 2026.
  20. ↑ "'Certain topics are temporarily restricted'". Meduza. 31 January 2025. Retrieved 12 June 2026.