Robbyant/lingbot-mapPublic

A feed-forward 3D foundation model for reconstructing scenes from streaming data

AI summary: A Geometric Context Transformer designed for high-performance streaming 3D reconstruction.

Stars
16.3K
+109 today
Forks
1.8K
Watchers
166
Open issues
51
Open PRs
18
Contributors
~3
Commits
110
Branches
1

PythonApache-2.0Created Apr 15, 2026Last push 14d ago+294 stars this week+622 this month

Star history

since Apr 12, 2026
05K10K15KApr 2026May 2026Jun 2026Aug 2026
16.3K stars as of Aug 7, 2026, tracked back to Apr 12, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulMonWedFri2025-08-03: 0 commits2025-08-04: 0 commits2025-08-05: 0 commits2025-08-06: 0 commits2025-08-07: 0 commits2025-08-08: 0 commits2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 18 commits2026-04-17: 6 commits2026-04-18: 6 commits2026-04-19: 0 commits2026-04-20: 6 commits2026-04-21: 9 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 4 commits2026-04-25: 0 commits2026-04-26: 2 commits2026-04-27: 3 commits2026-04-28: 12 commits2026-04-29: 2 commits2026-04-30: 3 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 1 commit2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 6 commits2026-05-26: 2 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 1 commit2026-06-01: 0 commits2026-06-02: 1 commit2026-06-03: 1 commit2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 1 commit2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 2 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 9 commits2026-07-03: 3 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 1 commit2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 3 commits2026-07-13: 1 commit2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 5 commits2026-07-21: 1 commit2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 1 commit2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits
110 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    16,295 stars

  • Breakout launch

    16,295 stars in 114 days

  • Permissive license

    Apache-2.0

  • Repeat trending

    10 trending appearances

What lingbot-map does

Lingbot-map is an advanced computer vision project that introduces a Geometric Context Transformer optimized for streaming 3D reconstruction tasks. It processes continuous streams of depth or point-cloud data to build accurate 3D maps of environments in real-time. By leveraging transformer architectures adapted for spatial data, it intelligently infers missing geometric context and refines the alignment of incoming data frames. This approach significantly reduces drift and improves map fidelity compared to traditional SLAM algorithms. It is designed for integration into robotics, autonomous navigation systems, and augmented reality applications requiring fast spatial awareness.

This project is built for robotics engineers, computer vision researchers, and developers working on spatial computing. It requires deep knowledge of 3D mathematics and PyTorch.

  • Streaming Reconstruction: Processes incoming 3D sensor data in real-time to build continuously updating environmental maps.
  • Geometric Context Transformers: Utilizes adapted transformer networks to infer complex spatial relationships and correct map drift.
  • High-Fidelity Alignment: Intelligently aligns incoming data frames by predicting missing geometry, resulting in smoother 3D models.
  • ROS Integration: Provides ready-to-use nodes for easy integration into the Robot Operating System ecosystem.
  • Optimized Inference: Designed to run efficiently on edge hardware, making it suitable for deployment on mobile robots.

Where teams use it

Autonomous Navigation

Drones or ground robots using the streaming map to navigate complex, dynamic environments without pre-existing maps.

Augmented Reality Mapping

AR headsets processing spatial data to build a real-time mesh of a room for placing stable virtual objects.

Industrial Inspection

Automated systems generating high-fidelity 3D models of infrastructure continuously as they move through a facility.

SLAM Research

Computer vision researchers using the geometric transformer architecture as a baseline for new 3D mapping algorithms.

Getting started: pip install -r requirements.txt && python train.py

README

main branch

LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction

Robbyant Team

Paper PDF Project HuggingFace ModelScope License

teaser.mp4

🗺️ Meet LingBot-Map! We've built a feed-forward 3D foundation model for streaming 3D reconstruction! 🏗️🌍

LingBot-Map has focused on:

  • Geometric Context Transformer: Architecturally unifies coordinate grounding, dense geometric cues, and long-range drift correction within a single streaming framework through anchor context, pose-reference window, and trajectory memory.
  • High-Efficiency Streaming Inference: A feed-forward architecture with paged KV cache attention, enabling stable inference at ~20 FPS on 518×378 resolution over long sequences exceeding 10,000 frames.
  • State-of-the-Art Reconstruction: Superior performance on diverse benchmarks compared to both existing streaming and iterative optimization-based approaches.

📑 Table of Contents

Click to expand

📰 News

  • 2026-06-28 — Fixed an SDPA KV cache bug. The SDPA backend now performs better on long sequences. We still recommend the FlashInfer backend for the best performance.
  • 2026-05-25 — 📊 Evaluation benchmark released. We released the evaluation scripts for KITTI and Oxford Spires — see benchmark/ for the pipeline, and run preprocess/oxford.py to prepare Oxford Spires data before evaluation.
  • 2026-04-29 — 📹 Long-video demo released. We released a very-long-video example (~25 000 frames, 13-minute indoor walkthrough) rendered with the offline pipeline — see Worked Example for the command, flag rationale, and rendered output.
  • 2026-04-27 — 🚀 LingBot-Map accelerated. Pull the latest main and run python demo.py --compile ... or python gct_profile.py --backend flashinfer --dtype bf16 --compile to verify on your hardware.
  • 2026-04-24 — Fixed a FlashInfer KV cache bug where --keyframe_interval > 1 silently cached non-keyframes. You should now see better pose and reconstruction quality when running with more than 320 frames.

📋 TODO

  • ✅ Release evaluation benchmark
    • ✅ Oxford Spires dataset
    • ✅ KITTI dataset
    • ✅ VBR dataset
    • ✅ Droid-W dataset
    • ✅ TUM-D dataset
    • ✅ 7-scenes dataset
    • ✅ ETH3D dataset
    • ✅ Tanks and Temples dataset
    • ✅ NRGBD dataset
  • ✅ Release demo scripts

⚙️ Installation

1. Create conda environment

conda create -n lingbot-map python=3.10 -y
conda activate lingbot-map

2. Install PyTorch (CUDA 12.8)

pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128

PyTorch 2.8.0 is the recommended version because NVIDIA Kaolin (required by the batch rendering pipeline) has prebuilt wheels for torch-2.8.0_cu128. If you only need demo.py you may use a newer PyTorch, but the batch renderer then requires building Kaolin from source. For other CUDA versions, see PyTorch Get Started.

3. Install lingbot-map

pip install -e .

4. Install FlashInfer (recommended)

FlashInfer provides paged KV cache attention for efficient streaming inference. It is a pure-Python package that JIT-compiles CUDA kernels on first use, so a single wheel works across CUDA/PyTorch versions:

pip install --index-url https://pypi.org/simple flashinfer-python

--index-url https://pypi.org/simple is only needed if your default pip index is an internal mirror that doesn't have flashinfer-python. (Optional) For faster first-use, you can additionally install a CUDA-specific JIT cache: pip install flashinfer-jit-cache -f https://flashinfer.ai/whl/cu128/flashinfer-jit-cache/. See FlashInfer installation for details. If FlashInfer is not installed, the model falls back to SDPA (PyTorch native attention) via --use_sdpa.

5. Visualization dependencies (optional)

pip install -e ".[vis]"

📦 Model Download

Model Name Huggingface Repository ModelScope Repository Description
lingbot-map-long robbyant/lingbot-map Robbyant/lingbot-map Better suited for long sequences and large scale scenes.
lingbot-map robbyant/lingbot-map Robbyant/lingbot-map Balanced checkpoint (used in paper, benchmark and offline demo) — trade off all-around performance across short and long sequences.
lingbot-map-stage1 robbyant/lingbot-map Robbyant/lingbot-map Stage-1 training checkpoint of lingbot-map — can be loaded into the VGGT model for bidirectional inference (c2w).

🚧 Coming soon: we're training an stronger model that supports longer sequences — stay tuned.

🚀 Quick Start

After installation, run your first scene with one command:

python demo.py --model_path /path/to/lingbot-map.pt \
    --image_folder example/courthouse --mask_sky

This launches an interactive viser viewer at http://localhost:8080. See Interactive Demo below for the full set of scenes and flags, or jump to Offline Rendering Pipeline for long-sequence batch rendering.

🎬 Interactive Demo (demo.py)

Run demo.py for interactive 3D visualization via a browser-based viser viewer (default http://localhost:8080).

Try the Example Scenes

We provide three example scenes in example/ that you can run out of the box:

# courthouse scene
python demo.py --model_path /path/to/lingbot-map.pt \
    --image_folder example/courthouse --mask_sky
output_pointcloud_side_by_side.mp4
# University scene
python demo.py --model_path /path/to/lingbot-map.pt \
    --image_folder example/university --mask_sky
output_pointcloud_side_by_side.mp4
# Loop scene (loop closure trajectory)
python demo.py --model_path /path/to/lingbot-map.pt \
    --image_folder example/loop
output_pointcloud_side_by_side.mp4

🎯 Featured: indoor walkthrough (~25 000 frames, 13 minutes)

Sequence is too long for the interactive viser viewer — this clip was rendered with the Offline Rendering Pipeline. See that section for the full command.

We will provide more examples in the follow-up.

Dynamic Demo (From Droid-W)

Dataset: Download the demo sequences from robbyant/lingbot-map-demo on Hugging Face.

Example run on the dynamic sequence from the dataset above (sky masking on, 4 camera optimization iterations, keyframe every 2 frames):

Run the dynamic sequence with sky masking, 4 camera optimization iterations, and an input stride of 2:

python demo.py \
    --image_folder /path/to/dynamic\
    --model_path ../../Lingbot-Map/lingbot-map.pt \
    --camera_num_iterations 4 \
    --mask_sky \
    --stride 2
output_pointcloud_side_by_side_under_10MB.mp4
image

Streaming with Keyframe Interval

Use --keyframe_interval to reduce KV cache memory by only keeping every N-th frame as a keyframe. Non-keyframe frames still produce predictions but are not stored in the cache. This is useful for long sequences which exceed 320 frames (We train with video RoPE on 320 views, so performance degrades when the KV cache stores more than 320 views. Using a keyframe strategy allows inference over longer sequences.). In demo.py, the keyframe interval is calculated automatically.

Note on inference range. Our method does not perform state resetting by default, so the maximum inference range is bounded by the longest distance seen during training on the dataset. Beyond that distance, state resetting becomes necessary. If you observe pose collapse, switch to windowed mode (--mode windowed) — in most cases tuning --keyframe_interval alone is enough and the rest of the windowed parameters can stay at their defaults.

Windowed Inference (for long sequences, >3000 frames)

python demo.py --model_path /path/to/lingbot-map.pt \
    --video_path video.mp4 --fps 10 \
    --mode windowed --window_size 128 --overlap_keyframes 16 --keyframe_interval 2 

Sky Masking

Sky masking uses an ONNX sky segmentation model to filter out sky points from the reconstructed point cloud, which improves visualization quality for outdoor scenes.

Setup:

# Install onnxruntime (required)
pip install onnxruntime        # CPU
# or
pip install onnxruntime-gpu    # GPU (faster for large image sets)

The sky segmentation model (skyseg.onnx) will be automatically downloaded from HuggingFace on first use.

Usage:

python demo.py --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ --mask_sky

Sky masks are cached in <image_folder>_sky_masks/ so subsequent runs skip regeneration. You can also specify a custom cache directory with --sky_mask_dir, or save side-by-side mask visualizations with --sky_mask_visualization_dir:

python demo.py --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ --mask_sky \
    --sky_mask_dir /path/to/cached_masks/ \
    --sky_mask_visualization_dir /path/to/mask_viz/

Visualization Options

Argument Default Description
--port 8080 Viser viewer port
--conf_threshold 1.5 Visibility threshold for filtering low-confidence points
--point_size 0.00001 Point cloud point size
--downsample_factor 10 Spatial downsampling for point cloud display

Performance & Memory

Without FlashInfer (SDPA fallback)

python demo.py --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ --use_sdpa

Running on Limited GPU Memory

If you run into out-of-memory issues, try one (or both) of the following:

  • --offload_to_cpu — offload per-frame predictions to CPU during inference (on by default; use --no-offload_to_cpu only if you have memory to spare).
  • --num_scale_frames 2 — reduce the number of bidirectional scale frames from the default 8 down to 2, which shrinks the activation peak of the initial scale phase.

Faster Inference

Lower the number of iterative refinement steps in the camera head to trade a small amount of pose accuracy for wall-clock speed:

python demo.py --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ --camera_num_iterations 1

--camera_num_iterations defaults to 4; setting it to 1 skips three refinement passes in the camera head (and shrinks its KV cache by 4×).

🎥 Offline Rendering Pipeline (demo_render/batch_demo.py)

Use this pipeline when your sequence is too long for the interactive viser viewer — for example, the indoor walkthrough featured above. demo_render/batch_demo.py is the all-in-one offline entry point: feed it a video or a folder of images and it will run model inference and produce a headless point-cloud flythrough MP4 in a single command. It shares the same PyTorch / FlashInfer / checkpoint stack as demo.py.

For those constrained by limited VRAM or GPU usage, you may also refer to the implementation at: https://github.com/ureeey/lingbot-map-rtx4060-8g/commit/eeee84a89cc97c1e39b736b46df4ee315275700b

Install (extends the main install)

1. Rendering Python dependencies

pip install -e ".[vis,render]"

render pulls in open3d>=0.19 and pyyaml (the core numpy<2 constraint comes from the base lingbot-map install). Sky masking in this pipeline uses onnxruntime-gpu for batched segmentation; install it if you don't already have the CPU onnxruntime:

pip install onnxruntime-gpu

2. Kaolin — matches the PyTorch 2.8.0 + CUDA 12.8 recommended above:

pip install --index-url https://pypi.org/simple \
    kaolin -f https://nvidia-kaolin.s3.us-east-2.amazonaws.com/torch-2.8.0_cu128.html

--index-url https://pypi.org/simple bypasses any internal mirror that might otherwise serve the PyPI placeholder wheel (which raises ImportError on import). NVIDIA Kaolin does not publish prebuilt wheels for PyTorch 2.9.x — if you're on 2.9 for other reasons, build Kaolin from source (pip install --no-build-isolation git+https://github.com/NVIDIAGameWorks/kaolin.git, needs local CUDA toolkit). For other torch/CUDA combinations see NVIDIA Kaolin installation.

3. ffmpeg

sudo apt install ffmpeg    # or: brew install ffmpeg

4. CUDA extensions (required before first run)

cd demo_render/render_cuda_ext && python setup.py build_ext --inplace && cd ../..

This builds voxel_morton_ext and frustum_cull_ext in place — both are imported by rgbd_render for GPU voxelization and frustum culling.

Worked Example — long indoor walkthrough (~25 000 frames, 13 minutes)

Dataset: Download the example video from robbyant/lingbot-map-demo on Hugging Face.

    python demo_render/batch_demo.py \
    --video_path /data/demo_videos/indoor_travel.MP4 \
    --output_folder /data/outputs/indoor_travel/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed --window_size 128 \
    --keyframe_interval 10 --overlap_keyframes 8 \
    --sky_mask_dir /data/outputs/sky_masks \
    --sky_mask_visualization_dir /data/outputs/sky_mask_viz \
    --camera_vis default --keyframes_only_points \
    --frame_tag --frame_tag_position top_right \
    --save_predictions
image

Flag-by-flag rationale:

Flag Why it's there
--mode windowed --window_size 128 Sliding-window inference is required once the sequence exceeds the ~320-frame RoPE training range; each window resets the KV cache. window_size counts KV-cache slots, not actual frames — the first num_scale_frames (=8) slots hold the scale frames and the remaining 128 − 8 = 120 slots hold keyframes. With keyframe_interval = 13, one window therefore covers 8 + 120 × 13 = 1568 actual frames.
--keyframe_interval 10 Cache only every 10th frame as a keyframe. Non-keyframes still emit per-frame predictions but don't grow the KV cache
--overlap_keyframes 8 Adjacent windows share 8 keyframes of context, resolved internally to max(num_scale_frames, 8 × keyframe_interval) = 8 × 13 = 104 actual frames of overlap. Recommended whenever keyframe_interval > 1, to keep cross-window pose alignment stable.
--config demo_render/config/indoor.yaml Seed render/scene/camera/overlay defaults from the indoor preset (short depth, tighter follow cam). Any CLI flag the user explicitly passes still overrides the YAML value.
--sky_mask_dir / --sky_mask_visualization_dir Persist sky masks and their side-by-side visualizations to disk so subsequent reruns reuse them instead of re-running ONNX segmentation. (The render pipeline only consumes them when sky masking is enabled — by the YAML preset or by --mask_sky.)
--camera_vis default Overlay the trajectory trail + recent-frame points on the rendered video.
--keyframes_only_points Only unproject keyframe depth into the point cloud; non-keyframes still contribute their pose to the trajectory/frustum overlay. Keeps the cloud sparse for very long sequences.
--frame_tag --frame_tag_position top_right Stamp a <i> / <N> Frames counter in the top-right corner of the MP4.
--save_predictions Persist per-frame NPZs alongside the MP4. Useful for inspection or for re-rendering with different camera/overlay settings later.

Replacing keyframe_interval = 10 with image_stride = 10 speeds up rendering. Then, uncomment the camera follow section in demo_render/config/indoor.yaml and set the birdeye's ranges to [2000, 2500] to reproduce the indoor fly-through effect shown in the demo:

image
_pointcloud_combined_10MB.mp4

Worked Example — outdoor drive scene

Dataset: Download the example video from robbyant/lingbot-map-demo on Hugging Face.

    python demo_render/batch_demo.py \
    --video_path /data/demo_videos/drive_frames.mp4 \
    --output_folder /data/outputs/drive/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/outdoor_drive.yaml \
    --mode windowed --window_size 128 \
    --max_non_keyframe_gap 100 --overlap_keyframes 8 \
    --image_stride 1 \
    --sky_mask_dir /data/outputs/sky_masks \
    --sky_mask_visualization_dir /data/outputs/sky_mask_viz \
    --camera_vis default --keyframes_only_points \
    --frame_tag --frame_tag_position top_right \
    --save_predictions
image

What differs from the indoor walkthrough above:

Flag Why it's there
--config demo_render/config/outdoor_drive.yaml Seed defaults from the outdoor preset: sky masking enabled, deeper render range (max_depth: 250), and a follow cam tuned for vehicle trajectories with a final birdeye reveal.
--image_stride 1 Use every video frame. Increase it to subsample long or high-FPS drive footage.
--max_non_keyframe_gap 100 Upper bound on consecutive non-keyframes before a keyframe is forced. Only active with flow-based keyframe selection (--flow_threshold > 0); in the default fixed-interval mode it has no effect.

The remaining flags (--mode windowed --window_size 128, --overlap_keyframes 8, sky-mask caching, overlays, --save_predictions) carry over unchanged from the indoor example — see the flag-by-flag table above.

Worked Example — LingBot-World scenes

Reconstruct videos generated by LingBot-World, our world model — the same pipeline works on generated footage out of the box.

Dataset: Download the example videos (lingbo_world_frames.mp4, lingbo_world2_frames.mp4) from robbyant/lingbot-map-demo on Hugging Face.

    python demo_render/batch_demo.py \
    --video_path /data/demo_videos/lingbo_world_frames.mp4 \
    --output_folder /data/outputs/lingbo_world/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/outdoor_drive.yaml \
    --mode windowed --window_size 128 \
    --max_non_keyframe_gap 100 --overlap_keyframes 8 \
    --image_stride 1 \
    --sky_mask_dir /data/outputs/sky_masks \
    --sky_mask_visualization_dir /data/outputs/sky_mask_viz \
    --camera_vis default --keyframes_only_points \
    --frame_tag --frame_tag_position top_right \
    --save_predictions

For the second clip, run the same command with --video_path /data/demo_videos/lingbo_world2_frames.mp4 --output_folder /data/outputs/lingbo_world2/ (and separate --sky_mask_dir / --sky_mask_visualization_dir folders if you want to keep the cached masks apart).

All flags are identical to the outdoor drive scene above — only the input video and output folder change. See the drive scene and indoor walkthrough tables for the flag-by-flag rationale.

image image

Camera Path (YAML)

The virtual camera path is described by the camera.segments list in the YAML preset passed via --config. Edit the YAML to design your own shot — no need to touch CLI flags.

Built-in presets live in demo_render/config/: default.yaml, indoor.yaml, outdoor_drive.yaml. Copy one and edit the camera: block.

YAML structure

camera:
  fov: 60.0          # camera field of view in degrees
  transition: 30     # frames blended between adjacent segments
  segments:
    - mode: follow            # chase cam following the input trajectory
      frames: [0, 1500]       # rendered-frame range this segment covers (-1 = end)
      back_offset: 0.3        # how far behind the input camera (fraction of scene scale)
      up_offset: 0.08         # vertical lift above the input camera
      look_offset: 0.4        # how far ahead the lookat target points
      smooth_window: 30       # trajectory smoothing window in frames
    - mode: birdeye           # rise up for a top-down reveal of the whole scene
      frames: [1500, 1800]
      reveal_height_mult: 2.5 # birdeye height = scene scale × this factor
    - mode: follow            # drop back into chase cam
      frames: [1800, -1]
      back_offset: 0.3
      up_offset: 0.08
      look_offset: 0.4

transition controls how many frames are blended between adjacent segments; frames: [0, -1] means "the whole sequence".

Available modes

mode Behavior Tunable fields
follow Chase cam tracks the input trajectory with smooth offsets. The most cinematic option for walkthroughs. back_offset, up_offset, look_offset, smooth_window, scale_frames
birdeye Top-down reveal of the whole scene. Useful for hero / overview shots. reveal_height_mult
static Fixed eye + lookat, auto-derived from the segment's start frame.
pivot Fixed eye, lookat sweeps along the trajectory.

Single-shot YAML examples

Pure follow (most common):

camera:
  fov: 60.0
  segments:
    - mode: follow
      frames: [0, -1]
      back_offset: 0.3
      up_offset: 0.08
      look_offset: 0.4
      smooth_window: 30

Full birdeye (good for overview / hero shots):

camera:
  fov: 60.0
  segments:
    - mode: birdeye
      frames: [0, -1]
      reveal_height_mult: 2.5

Follow with birdeye inserts: just list multiple segments in order under segments: — adjacent segments are interpolated using transition frames.

Caveat: when --config loads a YAML preset, passing any segment-shaping CLI flag (--camera_mode, --back_offset, --up_offset, --look_offset, --smooth_window, --follow_scale_frames, --birdeye_start, --birdeye_duration, --reveal_height_mult) discards the YAML's segments and rebuilds the camera path from those flags instead. To stay fully YAML-driven, don't pass any of them on the command line.

Output files

For a given output name (e.g. <scene> or <video_name>):

File Description
<name>_pointcloud.mp4 Rendered point-cloud flythrough
<name>_pointcloud_rgb.mp4 Original RGB frames encoded as video
<name>_pointcloud_config.yaml Full config snapshot of this run
batch_results.json Per-scene success / duration summary

📜 License

This project is released under the Apache License 2.0. See LICENSE file for details.

📖 Citation

@article{chen2026geometric,
  title={Geometric Context Transformer for Streaming 3D Reconstruction},
  author={Chen, Lin-Zhuo and Gao, Jian and Chen, Yihang and Cheng, Ka Leong and Sun, Yipengjing and Hu, Liangxiao and Xue, Nan and Zhu, Xing and Shen, Yujun and Yao, Yao and Xu, Yinghao},
  journal={arXiv preprint arXiv:2604.14141},
  year={2026}
}

✨ Acknowledgments

We thank Shangzhan Zhang, Jianyuan Wang, Yudong Jin, Christian Rupprecht, and Xun Cao for their helpful discussions and support.

This work builds upon several excellent open-source projects:


View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Code frequency

additions and deletions
+15.8K-15.8KWeek of 2026-04-12: +14,158 linesWeek of 2026-04-12: -946 linesWeek of 2026-04-19: +550 linesWeek of 2026-04-19: -107 linesWeek of 2026-04-26: +10,767 linesWeek of 2026-04-26: -645 linesWeek of 2026-05-03: +1 linesWeek of 2026-05-03: -2 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +15,788 linesWeek of 2026-05-24: -10 linesWeek of 2026-05-31: +2,168 linesWeek of 2026-05-31: -459 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +3 linesWeek of 2026-06-14: -3 linesWeek of 2026-06-21: +2 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +181 linesWeek of 2026-06-28: -94 linesWeek of 2026-07-05: +2 linesWeek of 2026-07-05: -2 linesWeek of 2026-07-12: +13 linesWeek of 2026-07-12: -6 linesWeek of 2026-07-19: +39 linesWeek of 2026-07-19: -35 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesApr 12, 2026Jul 26, 2026
+43.7K lines added, -2.3K removed over the last year.

Commits per week

last 52 weeks
300Week of 2025-08-03: 0 commitsWeek of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 30 commitsWeek of 2026-04-19: 19 commitsWeek of 2026-04-26: 22 commitsWeek of 2026-05-03: 1 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 8 commitsWeek of 2026-05-31: 3 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 1 commitsWeek of 2026-06-21: 2 commitsWeek of 2026-06-28: 12 commitsWeek of 2026-07-05: 1 commitsWeek of 2026-07-12: 4 commitsWeek of 2026-07-19: 7 commitsWeek of 2026-07-26: 0 commitsAug 3, 2025Jul 26, 2026
110 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 2 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 3 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 1 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 3 commitsMon 1:00 — 1 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 1 commitsMon 11:00 — 2 commitsMon 12:00 — 0 commitsMon 13:00 — 0 commitsMon 14:00 — 2 commitsMon 15:00 — 1 commitsMon 16:00 — 0 commitsMon 17:00 — 0 commitsMon 18:00 — 3 commitsMon 19:00 — 1 commitsMon 20:00 — 1 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 7 commitsTue 0:00 — 1 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 1 commitsTue 10:00 — 0 commitsTue 11:00 — 4 commitsTue 12:00 — 3 commitsTue 13:00 — 1 commitsTue 14:00 — 3 commitsTue 15:00 — 2 commitsTue 16:00 — 1 commitsTue 17:00 — 2 commitsTue 18:00 — 1 commitsTue 19:00 — 1 commitsTue 20:00 — 2 commitsTue 21:00 — 0 commitsTue 22:00 — 2 commitsTue 23:00 — 1 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 1 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 0 commitsWed 11:00 — 0 commitsWed 12:00 — 2 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 1 commitsWed 17:00 — 0 commitsWed 18:00 — 0 commitsWed 19:00 — 0 commitsWed 20:00 — 0 commitsWed 21:00 — 0 commitsWed 22:00 — 0 commitsWed 23:00 — 0 commitsThu 0:00 — 1 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 2 commitsThu 10:00 — 7 commitsThu 11:00 — 6 commitsThu 12:00 — 1 commitsThu 13:00 — 0 commitsThu 14:00 — 1 commitsThu 15:00 — 1 commitsThu 16:00 — 2 commitsThu 17:00 — 0 commitsThu 18:00 — 4 commitsThu 19:00 — 1 commitsThu 20:00 — 0 commitsThu 21:00 — 3 commitsThu 22:00 — 0 commitsThu 23:00 — 3 commitsFri 0:00 — 1 commitsFri 1:00 — 0 commitsFri 2:00 — 1 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 1 commitsFri 10:00 — 0 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 0 commitsFri 15:00 — 2 commitsFri 16:00 — 1 commitsFri 17:00 — 4 commitsFri 18:00 — 1 commitsFri 19:00 — 1 commitsFri 20:00 — 2 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 1 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 3 commitsSat 3:00 — 1 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 1 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 1 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Aug 1, 2026monthly#12+7,577
Jul 31, 2026monthly#12+7,577
Jul 30, 2026monthly#10+7,511
Jul 29, 2026monthly#9+7,904
Jul 28, 2026monthly#8+8,189
Jul 27, 2026monthly#7+8,326
Jul 19, 2026daily#16+5
Apr 21, 2026daily#21+87
Apr 19, 2026daily#22+178
Apr 18, 2026daily#10+245