LunarFM: A Multimodal Foundation Model and Shared Representation Space for the Lunar Surface
With renewed global interest in lunar in-situ resource utilization and sustained presence, high-precision characterization of the lunar surface has become increasingly critical. While orbital remote-sensing data volumes are massive, heterogeneous instrument observations, sparse labels, and domain-specific workflows have fragmented scientific analysis. This paper introduces LunarFM, a multimodal foundation model designed to learn a shared representation of the lunar surface from diverse orbital measurements. The model integrates observations from six instruments across three lunar missions, projecting 18 input channels into a unified embedding space. Experiments demonstrate its effectiveness for downstream tasks including similarity search, few-shot resource mapping, mineral abundance regression, and geological unit classification. The research team has simultaneously released a co-registered multimodal dataset spanning 70°S to 70°N, a pre-trained masked autoencoder, and a 768-dimensional joint-representation embedding dataset, with all code and data open-sourced.
Background and Context
As humanity approaches the era of sustained lunar presence and in-situ resource utilization (ISRU), the ability to efficiently analyze vast volumes of orbital remote-sensing data has emerged as a critical bottleneck for scientific exploration and resource development. While the accumulation of observational data has been substantial, the field faces significant challenges due to the heterogeneity of instruments across different missions, the extreme sparsity of labeled data, and the fragmentation of domain-specific workflows. Traditional scientific analysis has largely relied on custom modeling pipelines tailored to specific tasks, which hinders knowledge reuse and creates silos of information that are difficult to integrate. This fragmentation limits the scalability of lunar science, making it difficult to draw comprehensive conclusions from the diverse array of measurements available.
To address these systemic issues, researchers have introduced LunarFM, a pioneering multimodal foundation model designed to learn a shared representation of the lunar surface from diverse orbital measurements. Unlike previous approaches that focused on single modalities or isolated tasks, LunarFM leverages deep learning techniques to extract highly generalizable features of the lunar surface from a wide variety of orbital data sources. This model represents a paradigm shift in lunar data science, moving away from specialized, task-specific models toward a unified foundation model architecture. By creating a shared representation space, LunarFM aims to solve the complex problem of multi-source data fusion, providing a robust infrastructure for scientific discovery even in low-resource scenarios where labeled data is scarce.
Deep Analysis
From a technical perspective, LunarFM employs an advanced multimodal masked autoencoder architecture to achieve deep fusion and representation learning across different data modalities. The model integrates observations from six key instruments spanning three distinct lunar missions, capturing multidimensional physical information ranging from topography to spectral composition. Systematically, the architecture maps up to 18 input channels of high-dimensional, heterogeneous data into a unified 768-dimensional shared embedding space. This design allows observations from different sources and with different physical properties to be aligned and compared within a single semantic space, facilitating a more holistic understanding of the lunar surface.
The training strategy utilizes a masked autoencoder approach, which forces the model to learn the intrinsic correlations and complementary information between various physical quantities of the lunar surface. This self-supervised learning method enables the model to capture complex geological structures and material distribution patterns even in the absence of large amounts of annotated data. The resulting shared embedding space preserves the rich details of the original data while significantly reducing the data preprocessing costs for downstream tasks. This efficient conversion from raw observations to scientific semantics allows for rapid analysis and comparison of lunar features without the need for extensive manual feature engineering or domain-specific calibration for each new dataset.
Industry Impact
The release of LunarFM has profound implications for the lunar exploration community and the broader field of earth and planetary sciences. The research team has open-sourced a co-registered multimodal observation dataset covering latitudes from 70°S to 70°N, along with pre-trained model weights and a corresponding embedding dataset. This provides global researchers with a high-quality, machine-readable standard benchmark, significantly lowering the barrier to entry for lunar data research. By providing these resources, the team enables a wider range of scientists to leverage advanced deep learning techniques without needing to build complex data pipelines from scratch.
Furthermore, this foundation model paradigm is expected to drive a transition in lunar science from data-driven to knowledge-driven approaches. It accelerates the development of ISRU technologies by providing real-time, precise scientific basis for future lunar base site selection and resource extraction. The multimodal foundation model architecture and shared representation learning methods employed by LunarFM also offer a replicable technical path for remote sensing data analysis on Mars and other celestial bodies. This has the potential to spark an innovation wave in planetary science, promoting a unified and deepened cognitive framework for solar system bodies.
Outlook
Experimental evaluations conducted by the research team demonstrate the effectiveness and generalization capabilities of LunarFM across several challenging downstream tasks, including similarity search, few-shot resource mapping, mineral abundance regression, and geological unit classification. The results indicate that the shared embedding space extracted by LunarFM enables high-precision resource mapping and mineral composition estimation even under few-shot conditions, significantly outperforming traditional methods based on single modalities or handcrafted features. In geological unit classification tasks, the model exhibited strong semantic understanding capabilities, accurately identifying boundaries and features of different geological structures.
Ablation studies further confirmed the importance of multimodal fusion in enhancing representation robustness. Single-modality inputs showed a noticeable performance decline in complex geological regions, whereas multimodal joint representations effectively compensated for data missingness and noise interference. These key findings not only prove the superiority of LunarFM in scientific analysis but also highlight its immense potential in resource-oriented analysis. As the lunar exploration community continues to generate more data, LunarFM provides a scalable and adaptable framework that can integrate new instruments and missions, ensuring that the scientific value of orbital data is fully realized in the coming years of sustained lunar presence.