Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a21ff8cb4a3b3014

Jump to content

// Workers AI · dad joke modeWhy did the geospatial foundation model go to therapy? It had boundary issues.

From Wikipedia, the free encyclopedia


A geospatial foundation model (GFM), also referred to as an Earth observation foundation model, represents a paradigm shift in Earth system modeling by utilizing data-centric artificial intelligence (DCAI) to train on petabytes of structured and unstructured geoscientific data.[1] Unlike conventional, task-specific deep learning systems, GFMs use massive cross-disciplinary datasets to model the Earth's interactive and complex dynamics, offering highly flexible task specifications, multimodal inputs/outputs, and advanced geoscientific knowledge representation.[citation needed]

Background and history

[edit]

The emergence of data-driven Earth system sciences is underpinned by the massive expansion of big data in geosciences, which currently generates dozens of petabytes of information with hundreds of terabytes transmitted on a daily basis.[2] Traditional machine learning and deep learning methodologies have been successfully applied across varied domains, including geological event monitoring (such as volcanic eruptions and earthquakes)[3][4][5] resource mapping and natural resource exploration,[6][7] weather forecasting, climate modeling[8][9][10] and remote sensing analysis for urban planning or land cover classification.[11][12][13]

Despite these early successes, traditional artificial intelligence approaches face prominent challenges:

  • Data Dependency: Traditional models heavily rely on extensive, high-quality annotated datasets, which are often limited or incomplete due to the variable nature of Earth systems.[14]
  • Limited Generalization: The capacity of static AI models to generalize often degrades significantly when encountering novel or distinct geological environments outside their training context.[15][16]
  • Lack of Interpretability: Opaque model outputs challenge the transparency required for geoscientists to firmly establish trust in AI insights.[17][18]

To address these limitations, foundation models (FMs) have transitioned into the geosciences from computer vision and natural language processing.[19] FMs offer emergent properties such as zero-shot adaptability and contextual reasoning derived from massive neural network scale and self-supervised learning on vast datasets.[20][21] Early efforts to pioneer GFMs have spanned multiple specializations[22][23][24] though broad adoption remains constrained by multi source data complexity and the inherent constraints of contemporary AI architectures.[22][23][24][25]

Architecture and core mechanisms

[edit]

GFM architectures primarily leverage the modeling capacity of the Transformer architecture, using self-attention mechanisms to process long-range dependencies and contextual relationships,[26] mitigating the historical constraints of recurrent neural networks (RNNs) and convolutional neural networks (CNNs).[27][28]

Transformer full architecture

Core transformer frameworks

[edit]
  • Vanilla Transformers: Composed of encoder and decoder blocks featuring multihead self-attention (MHSA) and position-wise feed-forward layers, to capture contextual sequence vectors uniformly across text length.[29][30]
  • Vision Transformers (ViTs): Adapt the architecture to computer vision tasks by dividing input visual data into localized 16 by 16 patches, flattening them with linear projections, and combining them with positional encodings through an encoder framework.[30] Notable standard variations include TNT,[31] and PVT.[32]
  • Vision-Language Transformers: Fuse multimodal visual and textual streams via unified attention mechanisms to capture cross-modal associations for visual question answering (VQA) and image captioning, historically building on models like VisualBERT,[33] Uniter,[34] and OSCAR[35] which combine BERT text processing with Faster-RCNN object detection pipelines.[33][34][35]

Pretraining taxonomy

[edit]

GFMs acquire general-purpose weights and broad structural representations through expansive training methods:[36]

  • Supervised Pretraining: Utilizes large-scale labeled datasets to establish model pillars, historically scaled up through CNN backbones like AlexNet,[37] VGG,[38] and ResNet[39] over ImageNet or equivalent repositories[40] and through Transformer platforms like BERT,[41] RoBERTa,[42] and ALBERT.[43]
  • Self-Supervised Pretraining (SSL): Acts as the cornerstone for modern large-scale systems by using pretext tasks to derive implicit pseudo-labels from unlabeled inputs, minimizing human annotation demands.[44][45][46] SSL paradigms include:
    • Generative Pretraining: Captures underlying statistical parameters by reconstructing obscured or generated data fields, utilizing frameworks like Variational Autoencoders (VAEs),[47] Context Encoders for inpainting,[48] Masked Autoencoders (MAEs),[49] and Generative Adversarial Networks (GANs) such as BigGAN[50] and SRGAN.[51]
    • Contrastive Pretraining: Optimizes feature representations by bringing matching or augmented data pairs closer in embedding space while repelling dissimilar samples. Core contrastive workflows are anchored by negative sampling (e.g., SimCLR,[52] MoCo iterations[53][54]), data clustering (e.g., DeepCluster,[55] SwAV[56]), teacher-student knowledge distillation (e.g., BYOL,[57] DINO variants[58]), and cross-correlation matrix redundancy reduction (e.g., Barlow Twins).[59]
    • Predictive Pretraining: Relies on concrete pretext calculations such as relative spatial patch alignment,[60] geometric spatial transformations (e.g., RotNet[61]), or spectral color mapping (e.g., Image Colorization tasks[62]).
  • Hybrid Pretraining: Integrates complementary strategies—such as cross-modal contrastive training matched with autoregressive decoding or self-supervised generation paired with downstream oversight—exemplified by multimodal models like CLIP[63] and DALL-E.[64]

Adaptation strategies

[edit]

Pretrained foundational backbones are tuned for specific target geoscientific tasks or localized tracking using parameter-efficient frameworks:[65][66]

  • Fine-Tuning: Alters the complete weight configuration using task-specific datasets, often employing smaller learning rates to preserve pretrained global representations.
  • Prompt Tuning: Adapts models to downstream contexts without altering underlying weights. In visual-spatial settings, these are deployed as vision-driven prompts (e.g., VPT,[67] DePT,[68] PVIT[69]), language-driven text prompts (e.g., CoOp,[70] PLOT[71]), or dual vision-language prompts (e.g., UPT,[72] DPT,[73] TPT[74]).
  • Adapter Tuning: Adds compact, trainable parameters to a frozen base network. These can follow a sequential setup within forward architectures (e.g., Res-adapt,[75] DAN,[76] LST,[77] Conv-Adapter[78]), run parallel to traditional sublayers (e.g., ViT-Adapter,[79] PESF-KD,[80] AdaptMLP[81]), or blend architectures dynamically inside multihead attention systems as mix adapters (e.g., PATT,[82] ETT,[83] PALT,[84] VQT[85]).
  • Parameter Tuning: Modifies parameter groups via reparameterization paths, isolating localized weight adjustments (e.g., LoRA,[86] DyLoRA[87]), targeting structural bias updates (e.g., Bitfit,[88] AdapterBias[89]), or modulating weight and bias parts uniformly (e.g., SSF[90]).[86][87][90]

Notable models

[edit]

Advancements in geoscientific foundation models map across distinct modality systems, expanding past foundational language architectures like GPT variants,[91] BERT versions[42] and general text reasoning techniques.[92]

Large Language Models (LLMs) for geoscience

[edit]
  • GeoBERT: Developed by retraining standard BERT models over 20 million internal geological and geoscientific records to perform field-specific query answering and localized textual summarization.[93]
  • BERT-E: Built through targeted domain transfer learning from SciBERT to specialize in earth science keyword classification.[94]
  • K2: Pioneered as a dedicated geoscience-specific language foundation model utilizing the unique instruction-driven GeoSignal training dataset and the GeoBench verification collection.[95]
  • CnGeoPLM & GeoBERTSegmenter: Optimized for non-English records; CnGeoPLM handles geological entity and relationship extraction over a specialized Chinese text corpus (GeoCorpus),[96] while GeoBERTSegmenter resolves complex geological term word tokenization.[97]
  • OceanGPT: Formulated as a dedicated language engine for oceanography and marine sciences using the DoInstruct data method and measured via the OceanBench platform.[98]
  • GeoGalactica: A massive specialized 30-billion parameter geoscience model trained explicitly on the extensive GeoCorpus text resource.[99]
  • Geospatial Copilots and GIS Engines: Implementations like GPT4GEO, GEOGPT, BB-GeoGPT, and GeoLLM utilize text, geometry, and location engines (such as OpenStreetMap) to evaluate geospatial questions, measure socioeconomic livelihoods, and build realistic automation assistants.[19][100][101][102][103][104][105][106]

Large Vision Models (LVMs) for geoscience

[edit]
  • RingMo Series: Introduces a scalable remote sensing model framework driven by masked image modeling over 2 million spatial scenes.[11] Variants include plain ViT models adjusting rotated window dimensions,[12] spatiotemporal sequence trackers (RingMo-Sense),[107] and resource-optimized edge networks (RingMo-Lite).[108]
  • Segment Anything Model (SAM) Remote Sensing Adaptations: Adaptations and automated tracking pipelines mapping SAM to spatial datasets include RingMo-SAM,[109] large-scale segment annotation generation tools like SAMRS[110] and RSPrompter,[111] boundary-constrained frameworks,[112] and targeted asset identifiers like GeoSAM for mobility networks.[113]
  • Prithvi: An open-source spatial transformer backbone engineered jointly by NASA and IBM, pretrained using more than 1 terabyte of multi-spectral Harmonized Landsat Sentinel-2 (HLS) satellite observations.[114]
  • Atmospheric Vision Models: Foundational frameworks applied directly to weather systems and climate dynamics include Climax,[115] W-MAE,[116] FourCastNet,[117] GraphCast,[118] Fengwu,[119] and Pangu-Weather.[120]
  • SkySense: A multi-modal visual architecture showing state-of-the-art generalization across global earth image understanding benchmarks.[118][119][120][121]

Large Vision-Language Models (LVLMs) for geoscience

[edit]
  • GeoChat: Introduced as the first grounded remote sensing visual-language model delivering multi-task conversational features over coordinate inputs.[122]
  • EarthGPT & SkyeyeGPT: Multi-modal frameworks deployed for open-set visual scene analysis, multi-sensor imagery captioning, and conversational instruction response.[123][124]
  • RemoteCLIP & CLIP-RS: Adapt cross-modal text-image contrastive training paradigms for remote sensing, using unmanned aerial vehicle views and massive image-caption datasets.[125][126]
  • Lightweight and Multimodal RSVQA Systems: Architectures like LIT-4-RSVQA[127] and prompt-guided visual platforms[128] track multi-spectral queries while evaluating structural dataset biases.[129][130]

Foundational model agents

[edit]

Building on autonomous multi-module agent definitions (comprising perception pipelines, brain modules for storage and decision making, and tool-action vectors),[131][132] specialized geospatial software agents have emerged:

  • RS-ChatGPT: Implements a prompt-driven tool orchestration wrapper enabling ChatGPT to map, segment, and count remote sensing data fields by deploying external vision library scripts.[133]
  • Change-Agent: Fuses multi-level change tracking models with text logic to detect, count, and explain terrain alterations over temporal satellite spans.[134]
  • RS-Agent: Merges expert knowledge libraries with text and image processing systems to resolve open-ended operational queries in spatial engineering tasks.[135]

Applications

[edit]

GFMs are deployed across diverse observation vectors spanning space, air, terrestrial ground systems, and deep oceans:

Remote sensing, environmental monitoring, and land use

[edit]

The application of spatial vision platforms significantly enhances pixel classification over hyper-spectral and multi-spectral arrays, automating geographic land-use and land-cover (LULC) segmentation workflows without labor-intensive feature manual curation.[136][24] These architectures track long-term deforestation trends, model carbon biomass reserves, isolate vegetation changes, and manage building extraction footprints or structural urban modifications over time.[137][138]

Climate dynamics and weather forecasting

[edit]

GFMs function as unified multi-source processing centers, assimilating disjointed observation data from weather stations, high-resolution radars, and orbiting satellites to generate contextually unified semantic representations.[139] Deployed predictive networks downscale coarse global grid targets, measure regional drought extensions, track severe typhoon pathways, and enhance the tracking velocity of extreme precipitation events.[140][141][142]

Terrestrial geophysics and seismology

[edit]

In terrestrial ground settings, deep spatial architectures manage high-dimensional waveforms, historical subsurface geological profiles, and live sensor measurements. This assists in modeling intricate crustal deformations, updating geological hazard maps, mapping deep mineral or petrochemical reserves, and evaluating earthquake behaviors, seismic waveforms, or chaotic landslide hazards.[143][144][145][146][147]

Oceanography and marine exploration

[edit]

GFMs support interactive three-dimensional visualization networks mapping sub-surface, biological, and environmental observation streams. Deployed across intermediate and deep midocean zones (from 100 to 1,000 meters deep), these systems leverage intelligent sensor fusion to explore sparse marine coordinates, evaluate temperature fields, simulate fluid dynamics, and optimize marine conservation choices within Deep Blue AI platforms.[148][149]

References

[edit]
  1. Zhang, Hao; Xu, Jin-Jian; Cui, Hong-Wei; Li, Lin; Yang, Yaowen; Tang, Chao-Sheng; Boers, Niklas (December 2025). "When Geoscience Meets Foundation Models: Toward a general geoscience artificial intelligence system". IEEE Geoscience and Remote Sensing Magazine. 13 (4): 79–118. arXiv:2309.06799. Bibcode:2025IGRSM..13d..79Z. doi:10.1109/MGRS.2024.3496478.
  2. Reichstein, Markus; Camps-Valls, Gustau; Stevens, Bjorn; Jung, Martin; Denzler, Joachim; Carvalhais, Nuno (February 2019). "Deep learning and process understanding for data-driven Earth system science". Nature. 566 (7743): 195–204. Bibcode:2019Natur.566..195R. doi:10.1038/s41586-019-0912-1. hdl:21.11116/0000-0003-0B7B-8. PMID 30760912.
  3. Bergen, Karianne J.; Johnson, Paul A.; de Hoop, Maarten V.; Beroza, Gregory C. (22 March 2019). "Machine learning for data-driven discovery in solid Earth geoscience". Science. 363 (6433) eaau0323. Bibcode:2019Sci...363.0323B. doi:10.1126/science.aau0323. OSTI 1547596. PMID 30898903.
  4. Anantrasirichai, N.; Biggs, J.; Albino, F.; Bull, D. (September 2019). "A deep learning approach to detecting volcano deformation from satellite imagery using synthetic datasets". Remote Sensing of Environment. 230 111179. arXiv:1905.07286. Bibcode:2019RSEnv.23011179A. doi:10.1016/j.rse.2019.04.032.
  5. Hess, Philipp; Drüke, Markus; Petri, Stefan; Strnad, Felix M.; Boers, Niklas (3 October 2022). "Physically constrained generative adversarial networks for improving precipitation fields from Earth system models". Nature Machine Intelligence. 4 (10): 828–839. Bibcode:2022NatMI...4..828H. doi:10.1038/s42256-022-00540-1.
  6. Ham, Yoo-Geun; Kim, Jeong-Hwan; Luo, Jing-Jia (26 September 2019). "Deep learning for multi-year ENSO forecasts". Nature. 573 (7775): 568–572. Bibcode:2019Natur.573..568H. doi:10.1038/s41586-019-1559-7. PMID 31534218.
  7. Mitsui, Takahito; Boers, Niklas (July 2021). "Seasonal prediction of Indian summer monsoon onset with echo state networks". Environmental Research Letters. 16 (7): 074024. Bibcode:2021ERL....16g4024M. doi:10.1088/1748-9326/ac0acb.
  8. Weyn, Jonathan A.; Durran, Dale R.; Caruana, Rich (August 2019). "Can Machines Learn to Predict Weather? Using Deep Learning to Predict Gridded 500-hPa Geopotential Height From Historical Weather Data". Journal of Advances in Modeling Earth Systems. 11 (8): 2680–2693. Bibcode:2019JAMES..11.2680W. doi:10.1029/2019MS001705.
  9. Bi, Kaifeng; Xie, Lingxi; Zhang, Hengheng; Chen, Xin; Gu, Xiaotao; Tian, Qi (20 July 2023). "Accurate medium-range global weather forecasting with 3D neural networks". Nature. 619 (7970): 533–538. Bibcode:2023Natur.619..533B. doi:10.1038/s41586-023-06185-3. PMC 10356604. PMID 37407823.
  10. Zhang, Yuchen; Long, Mingsheng; Chen, Kaiyuan; Xing, Lanxiang; Jin, Ronghua; Jordan, Michael I.; Wang, Jianmin (20 July 2023). "Skilful nowcasting of extreme precipitation with NowcastNet". Nature. 619 (7970): 526–532. Bibcode:2023Natur.619..526Z. doi:10.1038/s41586-023-06184-4. PMC 10356617. PMID 37407824.
  11. 1 2 Sun, Xian; Wang, Peijin; Lu, Wanxuan; Zhu, Zicong; Lu, Xiaonan; He, Qibin; Li, Junxi; Rong, Xuee; Yang, Zhujun; Chang, Hao; He, Qinglin; Yang, Guang; Wang, Ruiping; Lu, Jiwen; Fu, Kun (2023). "RingMo: A Remote Sensing Foundation Model With Masked Image Modeling". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–22. Bibcode:2023ITGRS..6194732S. doi:10.1109/TGRS.2022.3194732.
  12. 1 2 Wang, Di; Zhang, Qiming; Xu, Yufei; Zhang, Jing; Du, Bo; Tao, Dacheng; Zhang, Liangpei (2023). "Advancing Plain Vision Transformer Toward Remote Sensing Foundation Model". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–15. Bibcode:2023ITGRS..6122818W. doi:10.1109/TGRS.2022.3222818.
  13. Ding, Lei; Zhu, Kun; Peng, Daifeng; Tang, Hao; Yang, Kuiwu; Bruzzone, Lorenzo (2024). "Adapting Segment Anything Model for Change Detection in VHR Remote Sensing Images". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–11. arXiv:2309.01429. Bibcode:2024ITGRS..6268168D. doi:10.1109/TGRS.2024.3368168.
  14. Karpatne, Anuj; Ebert-Uphoff, Imme; Ravela, Sai; Babaie, Hassan Ali; Kumar, Vipin (2019). "Machine Learning for the Geosciences: Challenges and Opportunities". IEEE Transactions on Knowledge and Data Engineering. 31 (8): 1544–1554. arXiv:1711.04708. Bibcode:2019ITKDE..31.1544K. doi:10.1109/TKDE.2018.2861006.
  15. Aslam, Rana Waqar; Shu, Hong; Javid, Kanwal; Pervaiz, Shazia; Mustafa, Farhan; Raza, Danish; Ahmed, Bilal; Quddoos, Abdul; Al-Ahmadi, Saad; Hatamleh, Wesam Atef (February 2024). "Wetland identification through remote sensing: Insights into wetness, greenness, turbidity, temperature, and changing landscapes". Big Data Research. 35 100416. Bibcode:2024BDR....3500416A. doi:10.1016/j.bdr.2023.100416.
  16. Chen, Keyan; Chen, Bowen; Liu, Chenyang; Li, Wenyuan; Zou, Zhengxia; Shi, Zhenwei (2024). "RSMamba: Remote Sensing Image Classification With State Space Model". IEEE Geoscience and Remote Sensing Letters. 21: 1–5. arXiv:2403.19654. Bibcode:2024IGRSL..2107111C. doi:10.1109/LGRS.2024.3407111.
  17. Toms, Benjamin A.; Barnes, Elizabeth A.; Ebert-Uphoff, Imme (September 2020). "Physically Interpretable Neural Networks for the Geosciences: Applications to Earth System Variability". Journal of Advances in Modeling Earth Systems. 12 (9) e2019MS002002. arXiv:1912.01752. Bibcode:2020JAMES..1202002T. doi:10.1029/2019MS002002.
  18. Shen, Chaopeng; Appling, Alison P.; Gentine, Pierre; Bandai, Toshiyuki; Gupta, Hoshin; Tartakovsky, Alexandre; Baity-Jesi, Marco; Fenicia, Fabrizio; Kifer, Daniel; Li, Li; Liu, Xiaofeng; Ren, Wei; Zheng, Yi; Harman, Ciaran J.; Clark, Martyn; Farthing, Matthew; Feng, Dapeng; Kumar, Praveen; Aboelyazeed, Doaa; Rahmani, Farshid; Song, Yalan; Beck, Hylke E.; Bindas, Tadd; Dwivedi, Dipankar; Fang, Kuai; Höge, Marvin; Rackauckas, Chris; Mohanty, Binayak; Roy, Tirthankar; Xu, Chonggang; Lawson, Kathryn (11 July 2023). "Differentiable modelling to unify machine learning and physical models for geosciences". Nature Reviews Earth & Environment. 4 (8): 552–567. arXiv:2301.04027. Bibcode:2023NRvEE...4..552S. doi:10.1038/s43017-023-00450-9. hdl:10150/672240.
  19. 1 2 Bommasani, Rishi; Hudson, Drew A.; Adeli, Ehsan; Altman, Russ; Arora, Simran; Arx, Sydney von; Bernstein, Michael S.; Bohg, Jeannette; Bosselut, Antoine (2021). "On the Opportunities and Risks of Foundation Models". arXiv:2108.07258 [cs.LG].
  20. Värtinen, Susanna; Hämäläinen, Perttu; Guckelsberger, Christian (March 2024). "Generating Role-Playing Game Quests With GPT Language Models". IEEE Transactions on Games. 16 (1): 127–139. Bibcode:2024ITGam..16..127V. doi:10.1109/TG.2022.3228480.
  21. Hu, Yushi; Hua, Hang; Yang, Zhengyuan; Shi, Weijia; Smith, Noah A.; Luo, Jiebo (2023). "PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3". 2023 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2951–2963. Bibcode:2023iccv.conf..277H. doi:10.1109/ICCV51070.2023.00277. ISBN 979-8-3503-0718-4.
  22. 1 2 Li, Xiang; Wen, Congcong; Hu, Yuan; Yuan, Zhenghang; Zhu, Xiao Xiang (June 2024). "Vision-Language Models in Remote Sensing: Current progress and future trends". IEEE Geoscience and Remote Sensing Magazine. 12 (2): 32–66. arXiv:2305.05726. Bibcode:2024IGRSM..12b..32L. doi:10.1109/MGRS.2024.3383473.
  23. 1 2 Fuller, Anthony; Millard, Koreen; Green, James R. (2022). "SatViT: Pretraining Transformers for Earth Observation". IEEE Geoscience and Remote Sensing Letters. 19: 1–5. Bibcode:2022IGRSL..1901489F. doi:10.1109/LGRS.2022.3201489.
  24. 1 2 3 Hong, Danfeng; Zhang, Bing; Li, Xuyang; Li, Yuxuan; Li, Chenyu; Yao, Jing; Yokoya, Naoto; Li, Hao; Ghamisi, Pedram; Jia, Xiuping; Plaza, Antonio; Gamba, Paolo; Benediktsson, Jon Atli; Chanussot, Jocelyn (2024). "SpectralGPT: Spectral Remote Sensing Foundation Model". IEEE Transactions on Pattern Analysis and Machine Intelligence. 46 (8): 5227–5244. Bibcode:2024ITPAM..46.5227H. doi:10.1109/TPAMI.2024.3362475. PMID 38568772.
  25. Cui, Jiequan; Zhong, Zhisheng; Tian, Zhuotao; Liu, Shu; Yu, Bei; Jia, Jiaya (2024). "Generalized Parametric Contrastive Learning". IEEE Transactions on Pattern Analysis and Machine Intelligence. 46 (12): 7463–7474. Bibcode:2024ITPAM..46.7463C. doi:10.1109/TPAMI.2023.3278694. PMID 37216259.
  26. Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; Uszkoreit, Jakob; Jones, Llion; Gomez, Aidan N.; Kaiser, Lukasz; Polosukhin, Illia (2017). "Attention Is All You Need". arXiv:1706.03762 [cs.CL].
  27. Salem, Fathi M. (2022). "Recurrent Neural Networks (RNN)". Recurrent Neural Networks: From Simple to Gated Architectures. Cham: Springer International. pp. 43–67. doi:10.1007/978-3-030-89929-5_3. ISBN 978-3-030-89928-8.
  28. Alzubaidi, Laith; Zhang, Jinglan; Humaidi, Amjad J.; Al-Dujaili, Ayad; Duan, Ye; Al-Shamma, Omran; Santamaría, J.; Fadhel, Mohammed A.; Al-Amidie, Muthana; Farhan, Laith (2021). "Review of deep learning: concepts, CNN architectures, challenges, applications, future directions". Journal of Big Data. 8 (1) 53. doi:10.1186/s40537-021-00444-8. PMC 8010506. PMID 33816053.
  29. Yang, Eric; Li, Matthew D; Raghavan, Shruti; Deng, Francis; Lang, Min; Succi, Marc D; Huang, Ambrose J; Kalpathy-Cramer, Jayashree (August 2023). "Transformer versus traditional natural language processing: how much data is enough for automated radiology report classification?". The British Journal of Radiology. 96 (1149) 20220769. doi:10.1259/bjr.20220769. PMC 10461267. PMID 37162253.
  30. 1 2 Khan, Salman; Naseer, Muzammal; Hayat, Munawar; Zamir, Syed Waqas; Khan, Fahad Shahbaz; Shah, Mubarak (2022). "Transformers in Vision: A Survey". ACM Computing Surveys. 54 (10s): 1–41. arXiv:2101.01169. doi:10.1145/3505244.
  31. Han, Kai; Xiao, An; Wu, Enhua; Guo, Jianyuan; Xu, Chunjing; Wang, Yunhe (2021). "Transformer in Transformer". arXiv:2103.00112 [cs.CV].
  32. Wang, Wenhai; Xie, Enze; Li, Xiang; Fan, Deng-Ping; Song, Kaitao; Liang, Ding; Lu, Tong; Luo, Ping; Shao, Ling (2021). "Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions". 2021 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 548–558. Bibcode:2021iccv.conf...61W. doi:10.1109/ICCV48922.2021.00061. ISBN 978-1-6654-2812-5.
  33. 1 2 Li, Liunian Harold; Yatskar, Mark; Yin, Da; Hsieh, Cho-Jui; Chang, Kai-Wei (2020). "What Does BERT with Vision Look At?". Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 5265–5275. doi:10.18653/v1/2020.acl-main.469.
  34. 1 2 Chen, Yen-Chun; Li, Linjie; Yu, Licheng; Kholy, Ahmed El; Ahmed, Faisal; Gan, Zhe; Cheng, Yu; Liu, Jingjing (2019). "UNITER: UNiversal Image-TExt Representation Learning". arXiv:1909.11740 [cs.CV].
  35. 1 2 Li, Xiujun; Yin, Xi; Li, Chunyuan; Zhang, Pengchuan; Hu, Xiaowei; Zhang, Lei; Wang, Lijuan; Hu, Houdong; Dong, Li (2020). "Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks". arXiv:2004.06165 [cs.CV].
  36. Liu, Xiao; Zhang, Fanjin; Hou, Zhenyu; Mian, Li; Wang, Zhaoyu; Zhang, Jing; Tang, Jie (2021). "Self-supervised Learning: Generative or Contrastive". IEEE Transactions on Knowledge and Data Engineering: 1. arXiv:2006.08218. doi:10.1109/TKDE.2021.3090866.
  37. Krizhevsky, Alex; Sutskever, Ilya; Hinton, Geoffrey E. (24 May 2017). "ImageNet classification with deep convolutional neural networks". Communications of the ACM. 60 (6): 84–90. doi:10.1145/3065386.
  38. Simonyan, Karen; Zisserman, Andrew (2014). "Very Deep Convolutional Networks for Large-Scale Image Recognition". arXiv:1409.1556 [cs.CV].
  39. He, Kaiming; Zhang, Xiangyu; Ren, Shaoqing; Sun, Jian (2016). "Deep Residual Learning for Image Recognition". 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 770–778. doi:10.1109/CVPR.2016.90. ISBN 978-1-4673-8851-1.
  40. Deng, Jia; Dong, Wei; Socher, Richard; Li, Li-Jia; Kai Li; Li Fei-Fei (2009-06-01). "ImageNet: A large-scale hierarchical image database". 2009 IEEE Conference on Computer Vision and Pattern Recognition. pp. 248–255. Bibcode:2009cvpr.conf...27D. doi:10.1109/CVPR.2009.5206848. ISBN 978-1-4244-3992-8.
  41. Devlin, Jacob; Chang, Ming-Wei; Lee, Kenton; Toutanova, Kristina (2018). "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding". arXiv:1810.04805 [cs.CL].
  42. 1 2 Liu, Yinhan; Ott, Myle; Goyal, Naman; Du, Jingfei; Joshi, Mandar; Chen, Danqi; Levy, Omer; Lewis, Mike; Zettlemoyer, Luke (2019). "RoBERTa: A Robustly Optimized BERT Pretraining Approach". arXiv:1907.11692 [cs.CL].
  43. Lan, Zhenzhong; Chen, Mingda; Goodman, Sebastian; Gimpel, Kevin; Sharma, Piyush; Soricut, Radu (2019). "ALBERT: A Lite BERT for Self-supervised Learning of Language Representations". arXiv:1909.11942 [cs.CL].
  44. Jing, Longlong; Tian, Yingli (November 2021). "Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey". IEEE Transactions on Pattern Analysis and Machine Intelligence. 43 (11): 4037–4058. Bibcode:2021ITPAM..43.4037J. doi:10.1109/TPAMI.2020.2992393. PMID 32386141.
  45. Chen, Xingyu; Zhao, Liye; Xu, Jiawen; Liu, Zhikang; Zhang, Chengxin; Xu, Luxiang; Guo, Ning (2024). "A Self-Supervised Contrastive Denoising Autoencoder-Based Noise Suppression Method for Micro Thrust Measurement Signals Processing". IEEE Transactions on Instrumentation and Measurement. 73: 1–17. Bibcode:2024ITIM...7341134C. doi:10.1109/TIM.2023.3341134.
  46. 2023 International Conference on Digital Image Computing: Techniques and Applications (DICTA). Port Macquarie, Australia: IEEE. 2023. Bibcode:2023dict.conf....... doi:10.1109/dicta60407.2023. ISBN 979-8-3503-8220-4.[page needed]
  47. Kingma, Diederik P.; Welling, Max (2013). "Auto-Encoding Variational Bayes". arXiv:1312.6114 [stat.ML].
  48. Pathak, Deepak; Krahenbuhl, Philipp; Donahue, Jeff; Darrell, Trevor; Efros, Alexei A. (2016). "Context Encoders: Feature Learning by Inpainting". 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 2536–2544. arXiv:1604.07379. Bibcode:2016cvpr.conf..278P. doi:10.1109/CVPR.2016.278. ISBN 978-1-4673-8851-1.
  49. He, Kaiming; Chen, Xinlei; Xie, Saining; Li, Yanghao; Dollar, Piotr; Girshick, Ross (2022). "Masked Autoencoders Are Scalable Vision Learners". 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 15979–15988. Bibcode:2022cvpr.conf.1553H. doi:10.1109/CVPR52688.2022.01553. ISBN 978-1-6654-6946-3.
  50. Brock, Andrew; Donahue, Jeff; Simonyan, Karen (2018). "Large Scale GAN Training for High Fidelity Natural Image Synthesis". arXiv:1809.11096 [cs.LG].
  51. Ledig, Christian; Theis, Lucas; Huszar, Ferenc; Caballero, Jose; Cunningham, Andrew; Acosta, Alejandro; Aitken, Andrew; Tejani, Alykhan; Totz, Johannes; Wang, Zehan; Shi, Wenzhe (2017). "Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network". 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 105–114. Bibcode:2017cvpr.conf...19L. doi:10.1109/CVPR.2017.19. ISBN 978-1-5386-0457-1.
  52. Chen, Ting; Kornblith, Simon; Norouzi, Mohammad; Hinton, Geoffrey (2020). "A Simple Framework for Contrastive Learning of Visual Representations". arXiv:2002.05709 [cs.LG].
  53. He, Kaiming; Fan, Haoqi; Wu, Yuxin; Xie, Saining; Girshick, Ross (2020). "Momentum Contrast for Unsupervised Visual Representation Learning". 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 9726–9735. Bibcode:2020cvpr.conf..975H. doi:10.1109/CVPR42600.2020.00975. ISBN 978-1-7281-7168-5.
  54. Chen, Xinlei; Xie, Saining; He, Kaiming (2021). "An Empirical Study of Training Self-Supervised Vision Transformers". 2021 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 9620–9629. Bibcode:2021iccv.conf..950C. doi:10.1109/ICCV48922.2021.00950. ISBN 978-1-6654-2812-5.
  55. Caron, Mathilde; Bojanowski, Piotr; Joulin, Armand; Douze, Matthijs (2018). "Deep Clustering for Unsupervised Learning of Visual Features". arXiv:1807.05520 [cs.CV].
  56. Caron, Mathilde; Misra, Ishan; Mairal, Julien; Goyal, Priya; Bojanowski, Piotr; Joulin, Armand (2020). "Unsupervised Learning of Visual Features by Contrasting Cluster Assignments". arXiv:2006.09882 [cs.CV].
  57. Grill, Jean-Bastien; Strub, Florian; Altché, Florent; Tallec, Corentin; Richemond, Pierre H.; Buchatskaya, Elena; Doersch, Carl; Pires, Bernardo Avila; Guo, Zhaohan Daniel (2020). "Bootstrap your own latent: A new approach to self-supervised Learning". arXiv:2006.07733 [cs.LG].
  58. Oquab, Maxime; Darcet, Timothée; Moutakanni, Théo; Vo, Huy; Szafraniec, Marc; Khalidov, Vasil; Fernandez, Pierre; Haziza, Daniel; Massa, Francisco (2023). "DINOv2: Learning Robust Visual Features without Supervision". arXiv:2304.07193 [cs.CV].
  59. Zbontar, Jure; Jing, Li; Misra, Ishan; LeCun, Yann; Deny, Stéphane (2021). "Barlow Twins: Self-Supervised Learning via Redundancy Reduction". arXiv:2103.03230 [cs.CV].
  60. Noroozi, Mehdi; Favaro, Paolo (2016). "Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles". arXiv:1603.09246 [cs.CV].
  61. Gidaris, Spyros; Singh, Praveer; Komodakis, Nikos (2018). "Unsupervised Representation Learning by Predicting Image Rotations". arXiv:1803.07728 [cs.CV].
  62. Zhang, Richard; Isola, Phillip; Efros, Alexei A. (2016). "Colorful Image Colorization". arXiv:1603.08511 [cs.CV].
  63. Radford, Alec; Kim, Jong Wook; Hallacy, Chris; Ramesh, Aditya; Goh, Gabriel; Agarwal, Sandhini; Sastry, Girish; Askell, Amanda; Mishkin, Pamela (2021). "Learning Transferable Visual Models From Natural Language Supervision". arXiv:2103.00020 [cs.CV].
  64. Ramesh, Aditya; Dhariwal, Prafulla; Nichol, Alex; Chu, Casey; Chen, Mark (2022). "Hierarchical Text-Conditional Image Generation with CLIP Latents". arXiv:2204.06125 [cs.CV].
  65. Lialin, Vladislav; Deshpande, Vijeta; Yao, Xiaowei; Rumshisky, Anna (2023). "Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning". arXiv:2303.15647 [cs.CL].
  66. Yu, Bruce X.B.; Chang, Jianlong; Wang, Haixin; Liu, Lingbo; Wang, Shijie; Wang, Zhiyu; Lin, Junfan; Xie, Lingxi; Li, Haojie; Lin, Zhouchen; Tian, Qi; Chen, Chang Wen (31 December 2024). "Visual Tuning". ACM Computing Surveys. 56 (12): 1–38. doi:10.1145/3657632.
  67. Jia, Menglin; Tang, Luming; Chen, Bor-Chun; Cardie, Claire; Belongie, Serge; Hariharan, Bharath; Lim, Ser-Nam (2022). "Visual Prompt Tuning". arXiv:2203.12119 [cs.CV].
  68. Gao, Yunhe; Shi, Xingjian; Zhu, Yi; Wang, Hao; Tang, Zhiqiang; Zhou, Xiong; Li, Mu; Metaxas, Dimitris N. (2022). "Visual Prompt Tuning for Test-time Domain Adaptation". arXiv:2210.04831 [cs.CV].
  69. Herzig, Roei; Abramovich, Ofir; Ben Avraham, Elad; Arbelle, Assaf; Karlinsky, Leonid; Shamir, Ariel; Darrell, Trevor; Globerson, Amir (2024). "PromptonomyViT: Multi-Task Prompt Learning Improves Video Transformers using Synthetic Scene Data". 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). pp. 6789–6801. Bibcode:2024wacv.conf..666H. doi:10.1109/WACV57701.2024.00666. ISBN 979-8-3503-1892-0.
  70. Zhou, Kaiyang; Yang, Jingkang; Loy, Chen Change; Liu, Ziwei (September 2022). "Learning to Prompt for Vision-Language Models". International Journal of Computer Vision. 130 (9): 2337–2348. arXiv:2109.01134. Bibcode:2022IJCV..130.2337Z. doi:10.1007/s11263-022-01653-1.
  71. Chen, Guangyi; Yao, Weiran; Song, Xiangchen; Li, Xinyue; Rao, Yongming; Zhang, Kun (2022). "PLOT: Prompt Learning with Optimal Transport for Vision-Language Models". arXiv:2210.01253 [cs.CV].
  72. Zang, Yuhang; Li, Wei; Zhou, Kaiyang; Huang, Chen; Loy, Chen Change (2022). "Unified Vision and Language Prompt Learning". arXiv:2210.07225 [cs.CV].
  73. Xing, Yinghui; Wu, Qirui; Cheng, De; Zhang, Shizhou; Liang, Guoqiang; Wang, Peng; Zhang, Yanning (2024). "Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model". IEEE Transactions on Multimedia. 26: 2056–2068. Bibcode:2024ITMm...26.2056X. doi:10.1109/TMM.2023.3291588.
  74. Shu, Manli; Nie, Weili; Huang, De-An; Yu, Zhiding; Goldstein, Tom; Anandkumar, Anima; Xiao, Chaowei (2022). "Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models". arXiv:2209.07511 [cs.CV].
  75. Rebuffi, Sylvestre-Alvise; Bilen, Hakan; Vedaldi, Andrea (2017). "Learning multiple visual domains with residual adapters". arXiv:1705.08045 [cs.CV].
  76. Rosenfeld, Amir; Tsotsos, John K. (March 2020). "Incremental Learning Through Deep Adaptation". IEEE Transactions on Pattern Analysis and Machine Intelligence. 42 (3): 651–663. arXiv:1705.04228. Bibcode:2020ITPAM..42..651R. doi:10.1109/TPAMI.2018.2884462. PMID 30507526.
  77. Sung, Yi-Lin; Cho, Jaemin; Bansal, Mohit (2022). "LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer Learning". arXiv:2206.06522 [cs.CL].
  78. Chen, Hao; Tao, Ran; Zhang, Han; Wang, Yidong; Li, Xiang; Ye, Wei; Wang, Jindong; Hu, Guosheng; Savvides, Marios (2024). "Conv-Adapter: Exploring Parameter Efficient Transfer Learning for ConvNets". 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 1551–1561. Bibcode:2024cvpr.conf..162C. doi:10.1109/CVPRW63382.2024.00162. ISBN 979-8-3503-6547-4.
  79. Chen, Zhe; Duan, Yuchen; Wang, Wenhai; He, Junjun; Lu, Tong; Dai, Jifeng; Qiao, Yu (2022). "Vision Transformer Adapter for Dense Predictions". arXiv:2205.08534 [cs.CV].
  80. Rao, Jun; Meng, Xv; Ding, Liang; Qi, Shuhan; Liu, Xuebo; Zhang, Min; Tao, Dacheng (2024). "Parameter-Efficient and Student-Friendly Knowledge Distillation". IEEE Transactions on Multimedia. 26: 4230–4241. Bibcode:2024ITMm...26.4230R. doi:10.1109/TMM.2023.3321480.
  81. Chen, Shoufa; Ge, Chongjian; Tong, Zhan; Wang, Jiangliu; Song, Yibing; Wang, Jue; Luo, Ping (2022). "AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition". arXiv:2205.13535 [cs.CV].
  82. Yu, Bruce X. B.; Chang, Jianlong; Liu, Lingbo; Tian, Qi; Chen, Chang Wen (2022). "Towards a Unified View on Visual Parameter-Efficient Transfer Learning". arXiv:2210.00788 [cs.CV].
  83. Xu, Chengming; Yang, Siqian; Wang, Yabiao; Wang, Zhanxiong; Fu, Yanwei; Xue, Xiangyang (2023). "Exploring Efficient Few-shot Adaptation for Vision Transformers". arXiv:2301.02419 [cs.CV].
  84. Wu, Jiarun; Chen, Qingliang (14 February 2022). "Pruning Adapters with Lottery Ticket". Algorithms. 15 (2): 63. doi:10.3390/a15020063.
  85. Tu, Cheng-Hao; Mai, Zheda; Chao, Wei-Lun (2023). "Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer Learning". 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 7725–7735. Bibcode:2023cvpr.conf..746T. doi:10.1109/CVPR52729.2023.00746. ISBN 979-8-3503-0129-8.
  86. 1 2 Hu, Edward J.; Shen, Yelong; Wallis, Phillip; Allen-Zhu, Zeyuan; Li, Yuanzhi; Wang, Shean; Wang, Lu; Chen, Weizhu (2021). "LoRA: Low-Rank Adaptation of Large Language Models". arXiv:2106.09685 [cs.CL].
  87. 1 2 Valipour, Mojtaba; Rezagholizadeh, Mehdi; Kobyzev, Ivan; Ghodsi, Ali (2022). "DyLoRA: Parameter Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation". arXiv:2210.07558 [cs.CL].
  88. Ben-Zaken, Elad; Ravfogel, Shauli; Goldberg, Yoav (2021). "BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models". arXiv:2106.10199 [cs.LG].
  89. Fu, Chin-Lun; Chen, Zih-Ching; Lee, Yun-Ru; Lee, Hung-yi (2022). "AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks". arXiv:2205.00305 [cs.CL].
  90. 1 2 Lian, Dongze; Zhou, Daquan; Feng, Jiashi; Wang, Xinchao (2022). "Scaling & Shifting Your Features: A New Baseline for Efficient Model Tuning". arXiv:2210.08823 [cs.CV].
  91. OpenAI; Achiam, Josh; Adler, Steven; Agarwal, Sandhini; Ahmad, Lama; Akkaya, Ilge; Aleman, Florencia Leoni; Almeida, Diogo; Altenschmidt, Janko (2023). "GPT-4 Technical Report". arXiv:2303.08774 [cs.CL].
  92. Raffel, Colin; Shazeer, Noam; Roberts, Adam; Lee, Katherine; Narang, Sharan; Matena, Michael; Zhou, Yanqi; Li, Wei; Liu, Peter J. (2019). "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer". arXiv:1910.10683 [cs.LG].
  93. Denli, Huseyin; Chughtai, Hassan A.; Hughes, Brian; Gistri, Robert; Xu, Peng (2021). "Geoscience Language Processing for Exploration". Abu Dhabi International Petroleum Exhibition & Conference D031S102R003. OnePetro. Bibcode:2021adip.conf07766D. doi:10.2118/207766-MS.
  94. Ramachandran, R.; Ramasubramanian, M.; Koirala, P.; Gurung, I.; Maskey, M. (2022). "Language Model for Earth Science: Exploring Potential Downstream Applications as well as Current Challenges". IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium. pp. 4015–4018. Bibcode:2022igar.conf..989R. doi:10.1109/IGARSS46834.2022.9883682. ISBN 978-1-6654-2792-0.
  95. Deng, Cheng; Zhang, Tianhang; He, Zhongmou; Xu, Yi; Chen, Qiyuan; Shi, Yuanyuan; Fu, Luoyi; Zhang, Weinan; Wang, Xinbing (2023). "K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization". arXiv:2306.05064 [cs.CL].
  96. Ma, Kai; Zheng, Shuai; Tian, Miao; Qiu, Qinjun; Tan, Yongjian; Hu, Xinxin; Li, HaiYan; Xie, Zhong (December 2023). "CnGeoPLM: Contextual knowledge selection and embedding with pretrained language representation model for the geoscience domain". Earth Science Informatics. 16 (4): 3629–3646. Bibcode:2023EScIn..16.3629M. doi:10.1007/s12145-023-01112-6.
  97. Wei, Dongqi; Liu, Zhihao; Xu, Dexin; Ma, Kai; Tao, Liufeng; Xie, Zhong; Qiu, Qinjun; Pan, Shengyong (October 2022). "GeoBERTSegmenter: Word Segmentation of Chinese Texts in the Geoscience Domain Using the Improved BERT Model". Earth and Space Science. 9 (10) e2022EA002511. Bibcode:2022E&SS....902511W. doi:10.1029/2022EA002511.
  98. Bi, Zhen; Zhang, Ningyu; Xue, Yida; Ou, Yixin; Ji, Daxiong; Zheng, Guozhou; Chen, Huajun (2023). "OceanGPT: A Large Language Model for Ocean Science Tasks". arXiv:2310.02031 [cs.CL].
  99. Lin, Zhouhan; Deng, Cheng; Zhou, Le; Zhang, Tianhang; Xu, Yi; Xu, Yutong; He, Zhongmou; Shi, Yuanyuan; Dai, Beiya (2023). "GeoGalactica: A Scientific Large Language Model in Geoscience". arXiv:2401.00434 [cs.CL].
  100. Roberts, Jonathan; Lüddecke, Timo; Das, Sowmen; Han, Kai; Albanie, Samuel (2023). "GPT4GEO: How a Language Model Sees the World's Geography". arXiv:2306.00020 [cs.CL].
  101. Ji, Yuhan; Gao, Song (2023). "Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations". arXiv:2307.03678 [cs.CL].
  102. Mooney, Peter; Cui, Wencong; Guan, Boyuan; Juhász, Levente (2023). "Towards Understanding the Geospatial Skills of ChatGPT: Taking a Geographic Information Systems (GIS) Exam". Proceedings of the 6th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery. pp. 85–94. doi:10.1145/3615886.3627745. ISBN 979-8-4007-0348-5.
  103. Zhang, Yifan; Wei, Cheng; Wu, Shangyou; He, Zhengting; Yu, Wenhao (2023). "GeoGPT: Understanding and Processing Geospatial Tasks through An Autonomous GPT". arXiv:2307.07930 [cs.CL].
  104. Zhang, Yifan; Wang, Zhiyun; He, Zhengting; Li, Jingxuan; Mai, Gengchen; Lin, Jianfeng; Wei, Cheng; Yu, Wenhao (September 2024). "BB-GeoGPT: A framework for learning a large language model for geographic information science". Information Processing & Management. 61 (5) 103808. doi:10.1016/j.ipm.2024.103808.
  105. Manvi, Rohin; Khanna, Samar; Mai, Gengchen; Burke, Marshall; Lobell, David; Ermon, Stefano (2023). "GeoLLM: Extracting Geospatial Knowledge from Large Language Models". arXiv:2310.06213 [cs.CL].
  106. Singh, Simranjit; Fore, Michael; Stamoulis, Dimitrios (2024). "GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots". 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 585–594. doi:10.1109/CVPRW63382.2024.00063. ISBN 979-8-3503-6547-4.
  107. Yao, Fanglong; Lu, Wanxuan; Yang, Heming; Xu, Liangyu; Liu, Chenglong; Hu, Leiyi; Yu, Hongfeng; Liu, Nayu; Deng, Chubo; Tang, Deke; Chen, Changshuo; Yu, Jiaqi; Sun, Xian; Fu, Kun (2023). "RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution Disentangling". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–21. Bibcode:2023ITGRS..6116166Y. doi:10.1109/TGRS.2023.3316166.
  108. Wang, Yuelei; Zhang, Ting; Zhao, Liangjin; Hu, Lin; Wang, Zhechao; Niu, Ziqing; Cheng, Peirui; Chen, Kaiqiang; Zeng, Xuan; Wang, Zhirui; Wang, Hongqi; Sun, Xian (2024). "RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid Framework". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–20. arXiv:2309.09003. Bibcode:2024ITGRS..6260447W. doi:10.1109/TGRS.2024.3360447.
  109. Yan, Zhiyuan; Li, Junxi; Li, Xuexue; Zhou, Ruixue; Zhang, Wenkai; Feng, Yingchao; Diao, Wenhui; Fu, Kun; Sun, Xian (2023). "RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing Images". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–16. Bibcode:2023ITGRS..6132219Y. doi:10.1109/TGRS.2023.3332219.
  110. Wang, Di; Zhang, Jing; Du, Bo; Xu, Minqiang; Liu, Lin; Tao, Dacheng; Zhang, Liangpei (2023). "SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model". arXiv:2305.02034 [cs.CV].
  111. Chen, Keyan; Liu, Chenyang; Chen, Hao; Zhang, Haotian; Li, Wenyuan; Zou, Zhengxia; Shi, Zhenwei (2024). "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation Based on Visual Foundation Model". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–17. arXiv:2306.16269. Bibcode:2024ITGRS..6256074C. doi:10.1109/TGRS.2024.3356074.
  112. Ma, Xianping; Wu, Qianqian; Zhao, Xingyu; Zhang, Xiaokang; Pun, Man-On; Huang, Bo (2024). "SAM-Assisted Remote Sensing Imagery Semantic Segmentation with Object and Boundary Constraints". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–16. arXiv:2312.02464. Bibcode:2024ITGRS..6243420M. doi:10.1109/TGRS.2024.3443420.
  113. Sultan, Rafi Ibn; Li, Chengyin; Zhu, Hui; Khanduri, Prashant; Brocanelli, Marco; Zhu, Dongxiao (2023). "GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation". arXiv:2311.11319 [cs.CV].
  114. Jakubik, Johannes; Roy, Sujit; Phillips, C. E.; Fraccaro, Paolo; Godwin, Denys; Zadrozny, Bianca; Szwarcman, Daniela; Gomes, Carlos; Nyirjesy, Gabby (2023). "Foundation Models for Generalist Geospatial Artificial Intelligence". arXiv:2310.18660 [cs.CV].
  115. Nguyen, Tung; Brandstetter, Johannes; Kapoor, Ashish; Gupta, Jayesh K.; Grover, Aditya (2023). "ClimaX: A foundation model for weather and climate". arXiv:2301.10343 [cs.LG].
  116. Man, Xin; Zhang, Chenghong; Feng, Jin; Li, Changyu; Shao, Jie (2023). "W-MAE: Pre-trained weather model with masked autoencoder for multi-variable weather forecasting". arXiv:2304.08754 [cs.LG].
  117. Kurth, Thorsten; Subramanian, Shashank; Harrington, Peter; Pathak, Jaideep; Mardani, Morteza; Hall, David; Miele, Andrea; Kashinath, Karthik; Anandkumar, Animashree (2022). "FourCastNet: Accelerating Global High-Resolution Weather Forecasting using Adaptive Fourier Neural Operators". arXiv:2208.05419 [physics.ao-ph].
  118. 1 2 Lam, Remi; Sanchez-Gonzalez, Alvaro; Willson, Matthew; Wirnsberger, Peter; Fortunato, Meire; Alet, Ferran; Ravuri, Suman; Ewalds, Timo; Eaton-Rosen, Zach (2022). "GraphCast: Learning skillful medium-range global weather forecasting". arXiv:2212.12794 [cs.LG].
  119. 1 2 Chen, Kang; Han, Tao; Gong, Junchao; Bai, Lei; Ling, Fenghua; Luo, Jing-Jia; Chen, Xi; Ma, Leiming; Zhang, Tianning; Su, Rui; Ci, Yuanzheng; Li, Bin; Yang, Xiaokang; Ouyang, Wanli (2023). "FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead". Communications Earth & Environment. 6 (1): 518. arXiv:2304.02948. doi:10.1038/s43247-025-02502-y.
  120. 1 2 Bi, Kaifeng; Xie, Lingxi; Zhang, Hengheng; Chen, Xin; Gu, Xiaotao; Tian, Qi (20 July 2023). "Accurate medium-range global weather forecasting with 3D neural networks". Nature. 619 (7970): 533–538. Bibcode:2023Natur.619..533B. doi:10.1038/s41586-023-06185-3. PMC 10356604. PMID 37407823.
  121. Guo, Xin; Lao, Jiangwei; Dang, Bo; Zhang, Yingying; Yu, Lei; Ru, Lixiang; Zhong, Liheng; Huang, Ziyuan; Wu, Kang (2023). "SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery". arXiv:2312.10115 [cs.CV].
  122. Kuckreja, Kartik; Danish, Muhammad Sohail; Naseer, Muzammal; Das, Abhijit; Khan, Salman; Khan, Fahad Shahbaz (2023). "GeoChat: Grounded Large Vision-Language Model for Remote Sensing". arXiv:2311.15826 [cs.CV].
  123. Zhang, Wei; Cai, Miaoxin; Zhang, Tong; Zhuang, Yin; Mao, Xuerui (2024). "EarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–20. arXiv:2401.16822. Bibcode:2024ITGRS..6209624Z. doi:10.1109/TGRS.2024.3409624.
  124. Zhan, Yang; Xiong, Zhitong; Yuan, Yuan (March 2025). "SkyEyeGPT: Unifying remote sensing vision-language tasks via instruction tuning with large language model". ISPRS Journal of Photogrammetry and Remote Sensing. 221: 64–77. arXiv:2401.09712. Bibcode:2025JPRS..221...64Z. doi:10.1016/j.isprsjprs.2025.01.020.
  125. Liu, Fan; Chen, Delong; Guan, Zhangqingyun; Zhou, Xiaocong; Zhu, Jiale; Ye, Qiaolin; Fu, Liyong; Zhou, Jun (2024). "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–16. arXiv:2306.11029. Bibcode:2024ITGRS..6290838L. doi:10.1109/TGRS.2024.3390838.
  126. Yuan, Zhiqiang; Zhang, Wenkai; Tian, Changyuan; Rong, Xuee; Zhang, Zhengyuan; Wang, Hongqi; Fu, Kun; Sun, Xian (2022). "Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information". IEEE Transactions on Geoscience and Remote Sensing. 60: 1–16. arXiv:2204.09860. Bibcode:2022ITGRS..6063706Y. doi:10.1109/TGRS.2022.3163706.
  127. Hackel, Leonard; Clasen, Kai Norman; Ravanbakhsh, Mahdyar; Demir, Begüm (2023). "LIT-4-RSVQA: Lightweight Transformer-Based Visual Question Answering in Remote Sensing". IGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing Symposium. pp. 2231–2234. Bibcode:2023igar.conf..561H. doi:10.1109/IGARSS52108.2023.10281674. ISBN 979-8-3503-2010-7.
  128. Chappuis, Christel; Zermatten, Valérie; Lobry, Sylvain; Le Saux, Bertrand; Tuia, Devis (2022). "Prompt–RSVQA: Prompting visual context to a language model for Remote Sensing Visual Question Answering". 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 1371–1380. Bibcode:2022cvpr.conf..143C. doi:10.1109/CVPRW56347.2022.00143. ISBN 978-1-6654-8739-9.
  129. Chappuis, Christel; Mendez, Vincent; Walt, Eliot; Lobry, Sylvain; Le Saux, Bertrand; Tuia, Devis (2022). "Language Transformers for Remote Sensing Visual Question Answering". IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium. pp. 4855–4858. Bibcode:2022igar.conf.1189C. doi:10.1109/IGARSS46834.2022.9884036. ISBN 978-1-6654-2792-0.
  130. Chappuis, Christel; Walt, Eliot; Mendez, Vincent; Lobry, Sylvain; Saux, Bertrand Le; Tuia, Devis (2023). "The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation". arXiv:2311.16782 [cs.CV].
  131. Xi, Zhiheng; Chen, Wenxiang; Guo, Xin; He, Wei; Ding, Yiwen; Hong, Boyang; Zhang, Ming; Wang, Junzhe; Jin, Senjie (2023). "The Rise and Potential of Large Language Model Based Agents: A Survey". arXiv:2309.07864 [cs.AI].
  132. Wang, Lei; Ma, Chen; Feng, Xueyang; Zhang, Zeyu; Yang, Hao; Zhang, Jingsen; Chen, Zhiyuan; Tang, Jiakai; Chen, Xu; Lin, Yankai; Zhao, Wayne Xin; Wei, Zhewei; Wen, Jirong (December 2024). "A survey on large language model based autonomous agents". Frontiers of Computer Science. 18 (6) 186345. doi:10.1007/s11704-024-40231-1.
  133. Guo, Haonan; Su, Xin; Wu, Chen; Du, Bo; Zhang, Liangpei; Li, Deren (2024). "Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models". arXiv:2401.09083 [cs.CV].
  134. Liu, Chenyang; Chen, Keyan; Zhang, Haotian; Qi, Zipeng; Zou, Zhengxia; Shi, Zhenwei (2024). "Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and Analysis". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–16. doi:10.1109/TGRS.2024.3425815.
  135. Xu, Wenjia; Yu, Zijian; Mu, Boyang; Wei, Zhiwei; Zhang, Yuanben; Li, Guangzuo; Wang, Jiuniu; Peng, Mugen (2024). "RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent". arXiv:2406.07089 [cs.CV].
  136. Yan, Zhiyuan; Li, Junxi; Li, Xuexue; Zhou, Ruixue; Zhang, Wenkai; Feng, Yingchao; Diao, Wenhui; Fu, Kun; Sun, Xian (2023). "RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing Images". IEEE Transactions on Geoscience and Remote Sensing. 61: 1–16. Bibcode:2023ITGRS..6132219Y. doi:10.1109/TGRS.2023.3332219.
  137. Wang, Mingze; Su, Lili; Yan, Cilin; Xu, Sheng; Yuan, Pengcheng; Jiang, Xiaolong; Zhang, Baochang (2024). "RSBuilding: Toward General Remote Sensing Image Building Extraction and Change Detection With Foundation Model". IEEE Transactions on Geoscience and Remote Sensing. 62: 1–17. arXiv:2403.07564. Bibcode:2024ITGRS..62S9395W. doi:10.1109/TGRS.2024.3439395.
  138. Sun, Jialin; Yan, Shuai; Alexandridis, Thomas; Yao, Xiaochuang; Zhou, Han; Gao, Bingbo; Huang, Jianxi; Yang, Jianyu; Li, Ying (24 April 2024). "Enhancing Crop Mapping through Automated Sample Generation Based on Segment Anything Model with Medium-Resolution Satellite Imagery". Remote Sensing. 16 (9): 1505. Bibcode:2024RemS...16.1505S. doi:10.3390/rs16091505.
  139. Panigrahi, Akash; Verma, Sagar; Terris, Matthieu; Vakalopoulou, Maria (2023). "Have Foundational Models Seen Satellite Images?". IGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing Symposium. pp. 4998–5001. doi:10.1109/IGARSS52108.2023.10283274. ISBN 979-8-3503-2010-7.
  140. Hess, Philipp; Drüke, Markus; Petri, Stefan; Strnad, Felix M.; Boers, Niklas (3 October 2022). "Physically constrained generative adversarial networks for improving precipitation fields from Earth system models". Nature Machine Intelligence. 4 (10): 828–839. Bibcode:2022NatMI...4..828H. doi:10.1038/s42256-022-00540-1.
  141. Bi, Kaifeng; Xie, Lingxi; Zhang, Hengheng; Chen, Xin; Gu, Xiaotao; Tian, Qi (20 July 2023). "Accurate medium-range global weather forecasting with 3D neural networks". Nature. 619 (7970): 533–538. Bibcode:2023Natur.619..533B. doi:10.1038/s41586-023-06185-3. PMC 10356604. PMID 37407823.
  142. Zhang, Yuchen; Long, Mingsheng; Chen, Kaiyuan; Xing, Lanxiang; Jin, Ronghua; Jordan, Michael I.; Wang, Jianmin (20 July 2023). "Skilful nowcasting of extreme precipitation with NowcastNet". Nature. 619 (7970): 526–532. Bibcode:2023Natur.619..526Z. doi:10.1038/s41586-023-06184-4. PMC 10356617. PMID 37407824.
  143. Bergen, Karianne J.; Johnson, Paul A.; de Hoop, Maarten V.; Beroza, Gregory C. (22 March 2019). "Machine learning for data-driven discovery in solid Earth geoscience". Science. 363 (6433) eaau0323. Bibcode:2019Sci...363.0323B. doi:10.1126/science.aau0323. OSTI 1547596. PMID 30898903.
  144. Ham, Yoo-Geun; Kim, Jeong-Hwan; Luo, Jing-Jia (26 September 2019). "Deep learning for multi-year ENSO forecasts". Nature. 573 (7775): 568–572. Bibcode:2019Natur.573..568H. doi:10.1038/s41586-019-1559-7. PMID 31534218.
  145. Mitsui, Takahito; Boers, Niklas (July 2021). "Seasonal prediction of Indian summer monsoon onset with echo state networks". Environmental Research Letters. 16 (7): 074024. Bibcode:2021ERL....16g4024M. doi:10.1088/1748-9326/ac0acb.
  146. Li, Gensheng; Song, Xianzhi; Tian, Shouceng; Zhu, Zhaopeng (November 2022). "Intelligent Drilling and Completion: A Review". Engineering. 18: 33–48. Bibcode:2022Engin..18...33L. doi:10.1016/j.eng.2022.07.014.
  147. Lawley, Christopher J. M.; Gadd, Michael G.; Parsa, Mohammad; Lederer, Graham W.; Graham, Garth E.; Ford, Arianne (2023). "Applications of Natural Language Processing to Geoscience Text Data and Prospectivity Modeling". Natural Resources Research. 32 (4): 1503–1527. Bibcode:2023NRR....32.1503L. doi:10.1007/s11053-023-10216-1.
  148. Sonnewald, Maike; Lguensat, Redouane; Jones, Daniel C; Dueben, Peter D; Brajard, Julien; Balaji, V (July 2021). "Bridging observations, theory and numerical simulation of the ocean using machine learning". Environmental Research Letters. 16 (7): 073008. arXiv:2104.12506. Bibcode:2021ERL....16g3008S. doi:10.1088/1748-9326/ac0eb0.
  149. Ma, Chunyan; Li, Xin; Li, Yujie; Tian, Xinliang; Wang, Yichuan; Kim, Hyoungseop; Serikawa, Seiichi (2021). "Visual information processing for deep-sea visual monitoring system". Cognitive Robotics. 1: 3–11. doi:10.1016/j.cogr.2020.12.002.