Back
NASA and IBM release an open source model for lunar science
SiTech AI Team3 min read

NASA and IBM release an open source model for lunar science

NASA and IBM Research have released an open source foundation model trained on 17 years of lunar observations, with strong results for polar ice prediction and coarse scale crater detection.

Lunar observations become a reusable model

NASA and IBM Research, working with several academic institutions, have released the NASA-IBM Lunar Foundation Model. They describe it as one of the first open source foundation models for lunar science. It is pretrained on large volumes of unlabeled data and can be adapted to specific tasks with a small number of labeled examples, addressing a common problem in lunar research where observations are plentiful but labels are scarce.

The team trained the model from scratch using SomBench, which they call the largest co-registered multimodal lunar corpus to date. It contains nearly 2 million tile bundles across 11 modalities and two spatial scales. About 1 million high-resolution images came from the Narrow Angle Camera at roughly 1 meter per pixel, while just under 964,000 multispectral images came from the Wide Angle Camera at 100 meters per pixel. The dataset has more than 30 spatially aligned data layers from nine instruments and four missions. It was split geographically by map zones to prevent leakage between training, validation and test data.

Strongest gains in ice and crater prediction

The model is based on TerraMind, a multimodal Earth observation model, but it was trained from scratch rather than fine-tuned. For each tile, it receives illumination angles, sun position, tile extent and other imaging geometry as explicit context. It breaks lunar images, elevation data and imaging geometry into tokens, then learns their relationships by predicting masked portions. High-resolution and coarse imagery are learned together, while FlexiViT allows the trained model to handle different image patch sizes without retraining.

The largest improvement was in predicting polar ice deposits. According to IBM, the model reduced prediction error by up to 22 percent compared with the best baseline, SwinV2-B. In coarse-scale crater detection, it beat SwinV2-B by nearly 19 percent while using only half the training data. The models performed similarly in meter-scale crater detection and Irregular Mare Patch segmentation, with most differences falling within variance between training runs. LoRA kept pace with full fine-tuning across the tested tasks and performed better on crater detection.

Useful for analysis, with positioning limits

IBM Research Europe director Juan Bernabé-Moreno said the model connects observations across instruments and reveals patterns that can be difficult to identify in isolation. It is not suited to absolute geodetic positioning. In generation tests, latitude and longitude were sometimes off by dozens of degrees, while reconstructed elevation structures could have shifted absolute height values. The researchers present it as a reusable foundation for downstream analysis rather than a replacement for physical measurement instruments. Controlled experiments isolating each innovation are still pending, and some test datasets are small.

The model is publicly available on Hugging Face, its code is on GitHub, and it is integrated into the open source TerraTorch toolkit. NASA and IBM also released machine learning-ready pretraining datasets and benchmark collections. The model forms part of their AI for Science collaboration, under which they have worked on foundation models since early 2022.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.