Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a402102d2cc926ee

Jump to content

Talk:Multimodal learning

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 1 month ago by Alenoach in topic Challenges

Untitled

[edit]

I think stacks of Boltzmann machines are usually called Deep Belief Networks, not Deep Boltzmann Machines. The article also does not discuss stacks of autoencoders. 194.117.26.63 (talk) 16:17, 9 May 2016 (UTC)Reply

Why is this page basically entirely about Boltzmann machines? Enervation (talk) 01:01, 22 December 2023 (UTC)Reply

Good question. This article should be modernized. It would perhaps also be worth having a separate article on large multimodal models. Alenoach (talk) 21:06, 27 February 2024 (UTC)Reply

Potential renaming

[edit]

Should the article be renamed to something like "Multimodality (machine learning)"? It's a bit of a lengthier title, but it seems more precise, since the focus here is not just on the learning anymore (the article has been pretty significantly modified). And when i make a Google search on "Multimodal learning", a lot of the articles seem to rather be related to education and how to teach students. Alenoach (talk) 19:14, 26 May 2024 (UTC)Reply

The title "Multimodal machine learning" could be another option. Alenoach (talk) 19:18, 26 May 2024 (UTC)Reply

Add a Challenges section

[edit]

The current article discusses specific architectures and applications but lacks the challenges in multimodal learning. I suggest adding a short section on these challenges, perhaps before "Applications". I am a co-author of one source below, so I would prefer not to make the change directly. Could any editor consider whether the following content would be useful? Thanks for taking a look.

Challenges

[edit]

Multimodal learning involves challenges in representing, aligning and combining different forms of information. These have been categorised as representation, translation, alignment, fusion and co-learning,[1] with a later taxonomy covering representation, alignment, reasoning, generation, transference and quantification.[2] In practice, real-world applications face additional specific challenges, including missing modalities, differences in data formats and structures across modalities, aligning information across modalities, combining complementary information without unnecessary redundancy, and privacy risks such as re-identification when multiple data sources are combined.[3] Xyl007 (talk) 15:20, 21 August 2026 (UTC)Reply

Hi Xyl007. A "Challenges" section could indeed be an improvement and your sourcing looks fine, but it's challenging to understand, particularly the second sentence. I expect that most readers likely have basic ML knowledge but don't understand what "transference", "alignment", "translation" and "quantification" mean in this context. It could for example be clearer if instead of having a second sentence that presents two extra lists of concepts that partially overlap with the one of the first sentence, there was some explanation of one or two basic concepts in simple terms.
A section like "Challenges" is typically added near the end of the article, before or after the section "Applications". Feel free to add your section once you think it's ready, and thanks for your contribution. Alenoach (talk) 20:49, 21 August 2026 (UTC)Reply
Hi Alenoach, thanks for your suggestions. I have simplified the content to make it more readable, as below. I will add this section, but please feel free to edit it further.
Multimodal learning is challenging because information can be represented in different formats or types across modalities. Two common challenges are alignment, determining which parts of different modalities correspond to each other, and fusion, determining how information across modalities should be combined.[4] Real-world applications face further challenges specific to multimodal AI, including missing modalities when a data source is unavailable, selecting modalities that provide complementary rather than redundant information, and privacy risks from combining data sources, which may allow individuals to be re-identified even if they are anonymous in each source separately.[5] xxx007 (talk) 11:20, 23 August 2026 (UTC)Reply
Looks good, thanks for your effort. Alenoach (talk) 14:36, 23 August 2026 (UTC)Reply

References

  1. Baltrušaitis, Tadas; Ahuja, Chaitanya; Morency, Louis-Philippe (2019). "Multimodal Machine Learning: A Survey and Taxonomy". IEEE Transactions on Pattern Analysis and Machine Intelligence. 41 (2): 423–443. doi:10.1109/TPAMI.2018.2798607.
  2. Liang, Paul Pu; Zadeh, Amir; Morency, Louis-Philippe (2024). "Foundations & Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions". ACM Computing Surveys. 56 (10): 1–42. doi:10.1145/3656580.
  3. Liu, Xianyuan; Zhang, Jiayang; Zhou, Shuo; et al. (2025). "Towards deployment-centric multimodal AI beyond vision and language". Nature Machine Intelligence. 7 (10): 1612–1624. doi:10.1038/s42256-025-01116-5.
  4. Baltrušaitis, Tadas; Ahuja, Chaitanya; Morency, Louis-Philippe (2019). "Multimodal Machine Learning: A Survey and Taxonomy". IEEE Transactions on Pattern Analysis and Machine Intelligence. 41 (2): 423–443. doi:10.1109/TPAMI.2018.2798607.
  5. Liu, Xianyuan; Zhang, Jiayang; Zhou, Shuo; et al. (2025). "Towards deployment-centric multimodal AI beyond vision and language". Nature Machine Intelligence. 7 (10): 1612–1624. doi:10.1038/s42256-025-01116-5.