
NASA and IBM’s Lunar AI Model Learns Across Instruments, Not Just Images
NASA and IBM’s open lunar model combines instruments and spatial scales. Its benchmarks show useful transfer, not proof of accessible Moon ice.
NASA and IBM’s Lunar AI Model Learns Across Instruments, Not Just Images
The same lunar crater can be a sharp rim in an optical image, a depression in a terrain model, and a cold region in a thermal product. Those views are related, but they are not interchangeable. Combining them requires knowing where each measurement belongs, how much ground it represents, and which apparent features come from the surface rather than the way the instrument observed it.
That is the scientific problem behind the NASA-IBM Lunar Foundation Model, whose open-source release IBM and NASA announced on September 10, 2026. The project couples a reusable remote-sensing model with a unified lunar dataset. It is intended to support research into craters, volcanic features, and potential polar ice by learning relationships across instruments and spatial resolutions, rather than treating every scientific task as a separate image-classification project.
The distinction is consequential. This is not a chatbot that has read about the Moon. According to the official model card, it is a vision-transformer encoder–decoder trained from scratch on lunar observations using an adaptation of TerraMind’s masked-token approach. Its output can become a representation for detection, segmentation, or regression after task-specific adaptation.
The same model card places a firm limit on the excitement: ice-prospectivity predictions are not measured ice, generated fields are not scientific-grade replacements for instruments, and the model has not been validated for landing-site certification or hazard clearance. Those are not peripheral caveats. They define what this release can contribute to lunar science and what would still require a different kind of evidence.
Before a lunar model can learn, its maps have to agree
IBM’s release announcement describes a machine-learning-ready dataset with over 30 spatially aligned layers from nine instruments across four missions. The named missions include Lunar Reconnaissance Orbiter, GRAIL, and Japan’s SELENE/Kaguya. The technical report also identifies Lunar Prospector hydrogen data among the inputs. The headline counts describe the wider data collection, not 30 independent camera images with identical footprints.
The measurements carry different physical meanings. LROC supplies optical observations; LOLA contributes topographic information; Diviner supplies thermophysical products; Mini-RF contributes radar observations; Kaguya products add mineralogical context; and GRAIL characterizes gravity variations. As the report explains, these sources collectively describe morphology, composition, temperature-related behavior, and subsurface structure. Their native spatial resolutions differ substantially, so putting them into one folder would not make them one training sample.
SomBench, the companion dataset, organizes that problem around optical anchors. One family uses LROC Wide Angle Camera imagery at approximately 100 meters per pixel. Another uses Narrow Angle Camera imagery around one meter per pixel, paired with co-registered higher-resolution terrain products. The report describes attaching auxiliary products to the same geographic bounds and preserving missing observations instead of silently filling them with invented measurements.
Co-registration is therefore part of the scientific contribution, not clerical preparation before the interesting AI work. If a thermal feature and an optical rim are offset, a model may learn a relationship that belongs to the processing pipeline rather than the Moon. Aligning products creates the possibility of meaningful cross-modal learning. It does not erase the uncertainty or native resolution of the original instruments.
That last point is easy to lose in a colorful map. Resampling a coarse product onto a finer grid changes its representation, not the amount of information the instrument measured. A gravity value covering a broad region does not become a local boulder measurement because it appears beside meter-scale imagery. The model must use regional context as context, not as permission to claim that every channel directly resolves the same small feature.
Eleven modalities do not mean eleven identical pictures
The model card’s accounting explains how the larger dataset becomes the model’s inputs: nearly two million co-registered tile bundles span eleven pretraining modalities. Nine are dense, image-like modalities across the two resolution families. The other two contain optical acquisition metadata and static-map context. This is compatible with the announcement’s larger layer count because several source products contribute context rather than separate full-resolution image streams.
For the regional WAC family, dense inputs include visible reflectance, ultraviolet reflectance, topography, slope, and aspect. For the NAC family, they include panchromatic imagery and higher-resolution terrain, slope, and aspect. The optical and terrain products within each family describe corresponding ground. A training sample belongs to one family or the other; it does not place both NAC and WAC image families together in the same sample.
The static-map branch is particularly important for understanding what multi-instrument means here. According to the card, it encodes tile-footprint averages from products such as thermophysics, radar, mineralogy, gravity, and hydrogen abundance where available. An average can provide useful regional context, but it discards variation inside the footprint. The model is not receiving a fully detailed image from every instrument for every tile.
Coverage is also uneven. The model card says some products are near-global while others are limited to polar regions or particular instrument footprints. Its high-resolution NAC pretraining is restricted to selected photometry sites and frames with co-registered stereo terrain models. Those sites are globally distributed, but their distribution should not be confused with globally dense high-resolution coverage. A new research area may differ materially from the best-represented training terrain.
For a scientist adapting the checkpoint, these distinctions determine what an experiment is actually testing. A result obtained with a rich polar map stack is not automatically evidence that the model works equally well with a single optical image. Nor does successful regional crater detection establish performance on every small, shadowed crater in a local NAC scene. The inputs and spatial regime are part of the result.
TerraMind’s recipe, rebuilt for lunar lighting
IBM Research describes TerraMind as the architectural starting point, drawing on its ability to combine data types and resolutions. The technical report is more precise: the lunar model was trained from scratch, not fine-tuned from TerraMind’s Earth-observation checkpoint. The transfer is a modeling recipe, with lunar-specific extensions, rather than evidence that Earth-trained weights already understood lunar terrain.
During pretraining, different modalities are represented through their own input adapters and tokenizers, then processed by a shared transformer. The model receives selected input information and learns to predict sampled target tokens. Repeated across co-registered observations, this masked-token objective is intended to teach relationships within and between modalities. Terrain and reflectance can constrain one another without requiring a human to label every crater in the pretraining collection.
The first lunar-specific extension gives the model acquisition geometry explicitly. Illumination angles, solar context, and tile location describe conditions under which the image was acquired. The report’s rationale is straightforward: if appearance depends strongly on lighting and that lighting information is already recorded, forcing the model to rediscover it from pixels wastes capacity and risks confusing illumination with geology.
The second extension trains both resolution families in one mixed-batch process. A single set of model weights receives learning signals from regional WAC samples and local NAC samples. This is not a claim that low-resolution measurements have been transformed into new high-resolution observations. It is a way to expose the learned representation to both spatial regimes while preserving the fact that they are different kinds of samples.
The model also uses FlexiViT patch-embedding adaptation, which the card says permits fine-tuning at different working patch sizes without retraining the backbone from scratch. Modality-wise processing allows downstream users to change the available input mixture. These are useful forms of flexibility for research, but each new configuration still needs evaluation. Support for a different patch grid is not a guarantee that every scientific quantity remains equally recoverable on it.
This conceptual diagram follows the report’s separation between training samples, shared weights, and downstream scientific tasks. The two input families meet in training the same model, not by pretending that their image pixels describe identical spatial scales.
flowchart TD
W[WAC regional tile bundles] --> M[Mixed-batch lunar pretraining]
N[NAC local tile bundles] --> M
M --> E[Shared lunar representation]
E --> C[Adapted crater detection]
E --> V[Adapted volcanic segmentation]
E --> I[Adapted ice-prospectivity regression]
C --> R[Comparison with reference annotations]
V --> R
I --> P[Comparison with derived prospectivity map]
The ice target is a map of scientific expectations
The most important sentence for interpreting the ice result is in the model card’s intended-use boundary: the target is a knowledge-driven fuzzy-overlay prospectivity map, not measured ice. The report’s polar-volatiles section explains that the task uses a stack of eight map layers, including thermal constraints, illumination state, terrain properties, and spatial context related to permanently shadowed regions.
A prospectivity map expresses where a set of scientific criteria suggests ice might be more likely. Here, the reference map was itself constructed from the input layers using a knowledge-driven combination rule. Training a model to reproduce that map tests whether it can learn the relationship between those inputs and the derived target. It does not test whether drilling at the highest-scoring location would recover water.
IBM’s announcement reports up to a 22% reduction in root mean squared error against a SwinV2-B ImageNet baseline for this task. RMSE measures disagreement between continuous predictions and their reference values, with larger individual errors penalized more strongly. Lower RMSE is better agreement with the target map. It is not a percentage of ice deposits correctly discovered, and it cannot be translated into a probability that a specific site contains usable ice.
This makes the result useful in a more limited, defensible way. A model that efficiently reproduces a physically motivated prospectivity relationship could help researchers organize candidate regions, examine how input combinations affect predictions, or develop follow-up experiments. Its usefulness would depend on preserving the distinction between the learned proxy and the physical resource the proxy is intended to guide scientists toward.
That resource question includes properties outside the benchmark: whether ice is actually present, how it is distributed, how deeply it is buried, and whether it could be accessed under local surface conditions. Neither the reported RMSE nor a bright region on a prediction map establishes concentration, recoverability, or the feasibility of extraction. The announcement supports research into potential deposits, not a new inventory of confirmed supplies for a lunar base.
Crater detection changes when the scale changes
The report’s crater benchmarks make the regional-versus-local distinction unusually concrete. The WAC task draws on the Robbins catalog, a large manually annotated reference for lunar impact craters. The NAC task uses a separate hand-labeled collection of smaller-scale scenes. They are not simply the same test viewed through two zoom settings. Different imagery, annotation procedures, and terrain selections shape what successful detection means.
For the regional benchmark, crater rim polygons are converted into bounding boxes. That transformation is useful for evaluating object detectors, but it also changes the information represented by a label. A box can identify where a crater sits without describing every irregularity along its rim. A high detection score should therefore not be read as a guarantee of detailed rim geometry suitable for a different geological or engineering calculation.
At the local scale, the report describes manual crater labeling using NAC imagery with co-registered terrain information. It partitions the collection at site and study-area levels to reduce leakage. A researcher interested in a new area should care about that organization because the appearance of a crater depends on more than its approximate diameter. Lighting, surrounding texture, degradation, and nearby overlapping features can all affect the visual task the detector must solve.
There is a useful hypothetical decision here for a team building a crater catalog: should the model maximize candidate retrieval for expert review, or should it produce a conservative set of detections with fewer false positives? Those objectives need not select the same operating threshold. The published mAP result helps compare systems, but a cataloging workflow would also need to inspect the actual missed and extra detections, particularly in the size ranges relevant to its scientific question.
This matters because the report identifies crater catalogs as inputs to relative-age analysis and geological mapping. An error pattern concentrated in one size range or terrain type could matter more to such an application than a modest change in an aggregate detection score. A proposed evaluation should examine those patterns rather than assume that all improvements in mAP are equally useful to every downstream interpretation.
The model’s attraction is that researchers can start both regional and local experiments from a shared representation. That could reduce duplicated model-development work. The evidence still has to remain scale-specific: the regional benchmark supports a stronger reported advantage, while the meter-scale comparison is much closer. Treating those outcomes separately is not an admission that the model failed its purpose; it is how to identify where reuse is already convincing and where a specialist detector remains competitive.
Reading the benchmark table without turning it into one accuracy score
The official model card reports the following results. These are NASA-IBM authors’ evaluations, not independent replication. Values are means with standard deviations across five random seeds. The table selects the best reported lunar-model adaptation for each row; it does not describe a single universal fine-tuning configuration that wins every task.
| Evaluation | Metric | NASA-IBM LFM | Best listed baseline |
|---|---|---|---|
| WAC craters, 50% training data | mAP, higher is better | 0.2541 ± 0.0018, full fine-tuning | 0.2313 ± 0.0027, SwinV2-B |
| WAC craters, 100% training data | mAP, higher is better | 0.2581 ± 0.0017, LoRA | 0.2420 ± 0.0047, SwinV2-B |
| NAC craters, meter scale | mAP, higher is better | 0.1543 ± 0.0098, LoRA | 0.1552 ± 0.0086, SwinV2-B |
| Irregular mare patches | Foreground IoU, higher is better | 0.5709 ± 0.0114, frozen encoder | 0.5687 ± 0.0181, ConvNeXtV2-B |
| Polar ice prospectivity | RMSE, lower is better | 0.0293 ± 0.0013, full fine-tuning | 0.0377 ± 0.0004, SwinV2-B |
The first WAC row concerns label efficiency. With half the training data, the authors’ lunar model exceeds the SwinV2-B result shown for that data fraction. More tellingly, its score also exceeds the listed SwinV2-B score using the full training set. That supports the narrower claim that lunar pretraining can reduce the labeled-data requirement for this particular regional crater benchmark. It does not establish a universal reduction in labeling work across lunar science.
The full-data WAC row shows that the advantage remains when more labels are available, with LoRA providing the highest listed lunar-model result. Mean average precision, or mAP, evaluates detection performance under the benchmark’s matching and scoring rules. It is not simply the share of lunar craters found. A model can produce different trade-offs between missing a crater, falsely detecting one, and locating its boundary poorly, all of which deserve examination beyond the headline score.
The NAC row is not a clean lunar-model win. The listed SwinV2-B mean is slightly higher, and the model card explicitly advises treating the leading systems as comparable because the margins are smaller than the variation across runs. It also notes that part of this benchmark was annotated at a coarser image scale and appears blurrier. The result supports useful transfer to meter-scale work, not a claim that the new model dominates every crater detector.
The irregular-mare-patch row is similarly close. Foreground intersection over union measures overlap between predicted patch regions and the reference masks. The lunar model’s best listed mean edges the strongest baseline, but the authors again caution that the difference is smaller than seed spread. The scientifically relevant signal is that a pretrained representation remains competitive in a difficult, label-limited segmentation task, not that a tiny numerical lead settles the geological interpretation.
The ice row presents the clearest listed margin, but it uses a different metric and a different target. Its lower RMSE should be read as stronger approximation of the prospectivity map. Comparing that decimal directly with a crater mAP or patch IoU would have no scientific meaning. Each score describes a different question, different labels, and different errors.
IBM’s press release also uses broader percentage language, including a nearly 19% context-scale crater comparison and a 3% volcanic-feature comparison against SwinV2-B. The model card’s table uses specific data fractions, adaptation choices, and best-baseline selections. Those comparisons should not be silently substituted for one another. The explicit rows above are a firmer basis for choosing a research experiment than one percentage presented as overall lunar accuracy.
A polar scientist’s next observation, not a mining decision
Consider a hypothetical team comparing two shadowed polar regions for follow-up study. This is an illustrative research workflow, not an announced NASA mission plan. Starting from the published ice-prospectivity task, the team could examine the model’s predictions alongside the thermal, terrain, and shadow-related layers that produced them. The decision would be where more investigation is informative, not where extraction is ready to begin.
Suppose the model favors one region, but that ranking changes when a thermal layer is withheld or perturbed within a plausible uncertainty range. That sensitivity would be scientifically useful information. It could suggest that the ranking depends heavily on one assumption rather than on agreement among independent observations. This proposed experiment is different from asserting that the released model already provides calibrated uncertainty intervals; the card makes no such general promise.
The team should also trace the target map’s construction. If the prediction agrees with a prospectivity rule because both depend on the same thermal product, the agreement is not independent confirmation of ice. A stronger follow-up would seek observations that challenge or corroborate the hypothesis through another measurement path. The model may help formulate the question, but the evidentiary value comes from what the next observation adds.
There is a spatial decision as well. The report describes the ice benchmark at 240 meters per pixel. A promising region at that scale is not a traversable route through local hazards, much less a drilling coordinate. Moving from regional prospectivity to surface operations requires additional maps, physical analysis, and mission-specific verification. Keeping those stages separate prevents the visual precision of a raster from being mistaken for operational certainty.
A volcanic boundary can be wrong before the model sees it
Irregular mare patches offer a different test of the release. IBM Research describes their ages and origins as scientifically debated, while the technical report explains that their subtle boundaries make them difficult segmentation targets. It also reports annotation adjustments to address misalignment associated with NAC pointing uncertainty. The reference outline is therefore not an infallible line drawn directly by nature.
Imagine a second hypothetical research team using the model to expand a catalog of these features. The useful first decision is how to review disagreement between a predicted mask and a published polygon. A larger predicted region could mean the model found a plausible extension, but it could also reflect lighting, a registration error, or a learned preference for nearby texture. Neither the prediction nor the old polygon should win automatically.
The team could compare the feature across available views, inspect the registration, and ask geologists to evaluate the disputed boundary before adding it to the catalog. If the model consistently highlights a class of ambiguous margins, that could motivate a better annotation protocol. Such a workflow would turn model disagreement into a scientific work queue rather than conceal it inside an aggregate overlap score.
Even an improved outline would not date the feature by itself. Segmentation determines which pixels belong to a mapped region under the labeling scheme; it does not directly settle the thermal history that produced the landform. IBM’s release connects the model to volcanic-history research, but the logical chain still runs from mapping to geological interpretation and additional evidence. A cleaner boundary can improve that chain without completing it.
What the authors’ controls establish, and what they leave open
The technical report describes geographic partitioning designed to keep overlapping observations from crossing training, validation, and test boundaries. That matters because neighboring or repeated lunar images can otherwise make an evaluation look like generalization when the model has already encountered essentially the same ground. The WAC and NAC tracks share the partition logic so that different views of one location are not casually treated as independent examples.
The authors also compare against an architecture-matched randomly initialized model. On ice prospectivity, the model card explains that some advantage comes from native token-level multimodal processing, while lunar pretraining contributes further improvement. That is more informative than attributing the entire gain to the training corpus. Architecture and pretraining are distinct ingredients, and the control helps reveal their combined effect.
Other ingredients remain less isolated. The card says ablations separating acquisition-geometry tokenization and mixed-resolution training from lunar pretraining as a whole remain to be run. The report also discloses different optimizer recipes across backbone families. Consequently, these evaluations compare practical model-and-adaptation setups, not a perfectly isolated experiment in which every training choice except pretraining data is identical.
Some test sets are small: the report lists ten test image–mask pairs for irregular mare patches and twenty-five test tiles for ice prospectivity. Repeating training with different random seeds measures sensitivity to that training process; it does not create new independent terrain or fully characterize uncertainty across the Moon. New-region evaluation and independent scientific replication would add evidence that the release benchmarks alone cannot supply.
Open weights make the limitations inspectable
The official repository card lists an Apache-2.0 license, downloadable checkpoints and tokenizers, and a TerraTorch-based adaptation path. This gives researchers a concrete starting point for reproducing and changing the experiments. The practical advantage is not merely avoiding a proprietary inference interface; it is being able to inspect which inputs, training choices, and task definitions support a claimed result.
Adaptation still involves choices. The card recommends low-rank adaptation, or LoRA, as a sensible starting point based on the authors’ experiments. LoRA trains small adapter components while leaving the other encoder weights frozen; the task head is still trained. Full fine-tuning changes the encoder more broadly and retains advantages on some smaller benchmarks. A fully frozen encoder is not universally sufficient: the authors report that it performs poorly on crater detection despite producing the best listed irregular-mare-patch result. A research group should choose adaptation by its target task rather than turn parameter efficiency into a blanket rule.
The most striking limitation concerns generated geography. The card says the model does not maintain a geodetic reference frame: generated terrain may preserve local shape while shifting absolute elevation, and generated coordinates can be badly displaced. Its multimodal generation examples are qualitative probes of learned relationships, not calibrated replacement maps. A convincing-looking terrain reconstruction can therefore be scientifically interesting and still be unsuitable for navigation or absolute measurement.
IBM and NASA have released a way to ask more questions of observations already collected, with enough technical detail to challenge the answers. The next meaningful result need not be a spectacular new lunar discovery. It could be a reproduced crater result on unfamiliar terrain, a corrected volcanic outline, or a polar prediction that fails when confronted with a better measurement. Each would teach researchers where this shared representation helps them see the Moon more clearly—and where they still need another instrument.
Sudeep Devkota is the editor of ShShell.ai, covering AI research and the technical evidence behind new models and scientific applications.