Guilhem Carmouze

Research internship, AIST, Tsukuba, Japan. April to August 2026.

Creation of a 360-degree navigation dataset using 3D Gaussian Splatting

Four months of research, from an A* planner and six-view panoramas in simulation to ArtiFixer-360, a pipeline that turns a plain pinhole video into a 360° video, reported with the run that passed its quality gates and the run that failed them.

Role
Research intern, in my fourth year. Author of nav_3dgs_pano and KachakaNavigation, and of the 360° extension in artifixer-360-pipeline, a derivative of NVIDIA ArtiFixer.
Team
Computer Vision Research Team, Artificial Intelligence Research Center (AIRC)
Period
15 April to 21 August 2026
Organisation
AIST, National Institute of Advanced Industrial Science and Technology, Tsukuba, Japan
Stack
Python, PyTorch, 3DGRUT, Splatfacto, COLMAP, DISCOVERSE, MuJoCo, ArtiFixer (video diffusion), ROS 2 Humble, PBS and Singularity on the ABCI cluster, pytest
One frame of the raw 3DGRUT render of the reconstructed scene, projected to an equirectangular panorama, and the same frame after the first ArtiFixer3D+ run on that 154-frame clip. Most holes and splatting noise are removed; residual warping and duplicated structures remain, and regions the camera never saw are generated, not observed. This early run is neither the 117-frame reference run that passed the gates nor the later run that failed them.

In three lines

Problem
A visual navigation model needs 360° observations with poses and commands. A scene rebuilt from a plain video only holds what the camera saw: about 12% of the sphere per pose.
What I built
A simulation that plans, drives and renders 4096 × 2048 panoramas inside a 3D Gaussian scene; a ROS 2 interface prepared for the Kachaka robot; and ArtiFixer-360, which repairs 14 overlapping views jointly with a video diffusion model and distils them back into the 3D scene.
Result
Depth-aware synchronisation lowered cross-view depth error by 27% before distillation, and a 117-frame reference run passed its gates. The full 154-frame run failed them, which led to a geometry-first redesign.

Context

From April to August 2026, in my fourth year, I was a research intern in the Computer Vision Research Team of the Artificial Intelligence Research Center (AIRC) at AIST, in Tsukuba, Japan. The internship ran in English and ends in a 29-page report.

The goal: produce 360° equirectangular observations, with poses and commands, for robot visual navigation, from scenes represented with 3D Gaussian Splatting (3DGS).

The difficulty is coverage. ArtiFixer, the repair model, was trained on pinhole video. A panorama needs the whole sphere, while the source camera used during development sees 12.06% of it at one pose (a 93.72° × 60.93° field of view). Everything else has to be rendered from Gaussians that were never observed from there.

An equirectangular panorama of a house interior: hallway, mirror, doors and wooden floor, bent by the projection.
Output: one frame of the equirectangular 360° video. The input is a plain pinhole video; it is third-party footage and is not shown here.

What I built

  1. Phase 1, May

    Navigation and panoramic rendering in a supplied 3DGS scene

    A DISCOVERSE and MuJoCo simulation: a 5 cm occupancy grid, A* planning and waypoint following. At each waypoint six co-located pinhole views are re-projected to a 4096 × 2048 equirectangular panorama with a custom overlapped cubemap (96° faces) and feather blending. Every panorama is logged with its position and heading; the velocity commands are recorded by the navigation script, which renders no panorama.

    Simulation only: the pose is simulator ground truth, the base is moved kinematically and the scene was supplied as a .ply file. No 3DGS training happens in that repository.

    Repository: nav_3dgs_panoabout 4,000 lines of Python, sole author

  2. Phase 2, May

    Building the scene from a plain video

    A COLMAP and Splatfacto baseline, measured on a 70/30 split at PSNR 26.08 dB and SSIM 0.91 after 30,000 iterations, then 28.69 dB and 0.94 after 60,000.

  3. Phase 3, June

    Reconstruction against world generation, and a robot interface

    I compared Matrix-3D, HY-World 2.0 and ExploreGS with navigation-oriented criteria: 360° coverage, useful radius and MEt3R p90. In parallel I prepared a ROS 2 Humble deployment interface for a visual navigation model (NoMaD) on the Kachaka mobile robot: four nodes (image relay, model, command sender, robot executor), stale-frame rejection at 0.5 s, a velocity clamp at 0.2 m/s and 0.5 rad/s, a 20 Hz dead-man timer, a dry-run mode and a typed wrapper over kachaka-api (gRPC).

    On main, the NoMaD inference is not wired in, and no complete navigation run on the real Kachaka is claimed.

    Repository: KachakaNavigationabout 3,600 lines plus tests

  4. Phase 4, July

    Repairing incomplete renders

    I applied NVIDIA ArtiFixer, a video diffusion model, to panoramas and diagnosed why they break: the source camera sees about 12% of the sphere per pose, and repairing the cube faces independently raised the seam-failure rate from 0.34 to 0.83.

  5. Phase 5, August

    The ArtiFixer-360 pipeline

    Pinhole video, COLMAP, a 3DGRUT Gaussian scene, then a world-locked rig of 14 overlapping 110° views that follows the real camera path. The 14 streams are repaired jointly by the 14B video diffusion model, synchronised during denoising through a depth and occlusion-aware reprojection graph, distilled back into the 3D scene with the geometry locked, and rendered as an equirectangular 360° video.

    A derivative of NVIDIA nv-tlabs/ArtiFixer, under Apache-2.0.

    Repository: artifixer-360-pipelinemy delta over upstream: +23,602 / −630 lines, 121 new files

ArtiFixer-360, from a plain video to a 360° video

  1. 01Pinhole video
  2. 02COLMAP poses
  3. 03Shared 3D Gaussian scene3DGRUT
  4. Render, repair, distil: the repaired views go back into the shared scene

    1. 04World-locked rig14 views, 110° field of view, real camera centres
    2. 05Joint repair14 streams, 77-frame windows, synchronised through a depth and occlusion-aware reprojection graph
    3. 06Geometry-locked distillationback into the 3D scene
  5. 07Frustum-to-ERP stitchingof the final renders
  6. 08360° ERP video
  7. 09Reference-free quality gatesdecide whether a run is published
Added or extended in my repositoryInput, output and upstream componentsCOLMAP preparation and the 3DGRUT scene come with upstream ArtiFixer, and the 14B diffusion model is NVIDIA’s. The rig, the joint 14-stream loop, the reprojection graph, the distillation controls, the stitcher and the quality gates were added or extended during the internship.

14

overlapping 110° views per pose: six on the horizon, four pitched up by 45° and four pitched down. They share the real camera centre, so two views differ by a pure rotation.

Two equirectangular maps. Left: number of rig views covering each direction, from 2 to 5, with the small field of view of the source camera outlined in yellow at the centre. Right: which of the 14 views owns each direction when stitching.
Coverage of the sphere by the 14-view rig, recomputed by a script in the repository: every direction is seen by at least 2 and at most 5 views, 3.27 on average. The yellow outline is what the source camera sees at one pose: 12.06% of the sphere.

What sits around the model

My delta over the upstream NVIDIA code is +23,602 / −630 lines across 145 files (121 new, 24 modified), including 32 new test modules (119 tests). It covers the trajectory and rig generators, the joint multi-view inference loop, the distillation controls, the frustum-to-panorama stitcher, and PBS and Singularity jobs for the 4-GPU nodes of the ABCI cluster. A companion patch to NVIDIA 3DGRUT-ArtiFixer adds +506 / −74 lines.

Because no ground-truth panorama exists, a run is judged by reference-free quality gates: cross-view depth overlap, optical-flow temporal warp, wrap-seam ratio, edge retention and a bitwise audit that the locked geometry did not move.

Results

−27%

Cross-view depth-overlap MAE of the repaired views, without and with depth-aware synchronisation: 0.0340 to 0.0247. Measured before distillation, on the repaired pseudo-views.

0.037→0.020

Optical-flow temporal warp MAE, raw renders against the output, on the 117-frame reference run: 14 views, 1,638 renders, about 20 minutes of wall-clock time on one 4-GPU node.

What was measuredValueHow to read it
Cross-view depth-overlap MAE of the repaired views, before distillation, without and with depth-aware synchronisation0.0340 → 0.0247Positive, limited scope: measured on the repaired views, not on the final panorama.
The same change after distillation, on the final panorama: temporal warp MAE, first run against the depth and loop variant0.0233 against 0.0258Negative: the gain did not clearly survive distillation.
117-frame reference run: temporal warp MAE, raw renders to output0.037 → 0.020Positive, with a caveat: any smoothing lowers this metric, so it is read together with edge retention.
Same run: edge strength kept in high-confidence regions (1 = fully kept)0.49 medianPart of the stability comes with softer detail.
Baseline before the pipeline: video, COLMAP, Splatfacto, 70/30 split26.08 dB / 0.91 at 30k, 28.69 dB / 0.94 at 60kContext for the reconstruction stage (PSNR / SSIM).
Values come from GPU runs made during the internship and are transcribed in the repository from its dated records. Lower is better, except for PSNR, SSIM and edge strength.
Two renders of the same hallway side by side: the first ArtiFixer3D+ run and the variant distilled with depth and loop constraints. Local structure differs, both still show distortions.
Left: first ArtiFixer3D+ run. Right: depth and loop distillation. The depth and loop branch changes local structure and appearance, but distortions remain: a qualitative diagnostic, not a claim of correct geometry.

What failed or is unfinished

154

frames in the full 14-direction run. It reached complete coverage and failed my own visual and temporal acceptance gates. That result led to a geometry-first redesign.

  • Repairing cube faces one by one made the seams worse. The seam-failure rate went from 0.34 to 0.83. The approach was rejected after measurement; the final pipeline repairs the 14 views jointly.

  • The depth gain did not survive distillation. Depth-aware synchronisation improves the agreement of the repaired views, but the current distillation does not preserve that gain in the final panorama.

  • No navigation run on the real robot. On main, the NoMaD inference is not wired into the ROS 2 interface, and no closed-loop run on the Kachaka is claimed.

  • Phase 1 lives in simulation. The pose is simulator ground truth, the base moves kinematically and the 3DGS scene was supplied, not trained there.

  • The May stitching code mirrored the panoramas. They came out mirrored left to right. I found and fixed it in October 2026, with synthetic tests and AI assistance; the fix is checked against MuJoCo’s own renderer and synthetic rooms, not confirmed on a render of the lab scene.

  • Not a solved problem. The method does not guarantee panoramas without visible seams or with correct geometry. The evidence is limited to two indoor clips and to reference-free metrics.

Credits

  • ArtiFixer-360 is a derivative of NVIDIA’s ArtiFixer (nv-tlabs/ArtiFixer, Apache-2.0). The diffusion model, the base inference code and 3DGRUT are NVIDIA’s work; my modifications are itemised file by file in the repository.

  • The 3DGS scene of phase 1 was supplied by the lab. The simulation runs on DISCOVERSE and MuJoCo.

  • The two indoor clips behind the results are third-party video that I did not film. The footage itself is not shown: every picture of those rooms on this page is a 3D Gaussian render or a model output derived from it. The repository does not record the source or the licence of either clip.

  • The GPU runs were made on 4-GPU nodes of the ABCI cluster.

  • AI assistance: as declared in the appendix of my report, AI tools were used for research, for writing code and for spell-checking, not for running the experiments or producing the results.