3D Gaussian Splatting
A photo gives me one view of a place. What interests me about 3D Gaussian Splatting is being able to move the camera after the photos have been taken.
The goal is novel-view synthesis: use photographs from several angles to render the scene from another viewpoint. The 2023 paper by Kerbl, Kopanas, Leimkühler, and Drettakis does this with a collection of soft, overlapping 3D shapes.
What is a Gaussian here?
Think of a small cloud of colour, strongest at its centre and fading towards its edges. Stretch it, flatten it, rotate it. We can fit these shapes around the structures in a scene.
A Gaussian stores a position, a shape, an opacity, and colour information. Its shape comes from a scale along each of three axes and a rotation. The colour can vary with viewing direction; the original method represents this with spherical harmonics, functions defined over directions.
The falloff is easier to see in one dimension:
Here controls the width. These curves have a peak of 1; they are not normalized probability densities.
In three dimensions, let be the offset from the centre. The same idea becomes:
where is the centre and is the covariance matrix, which describes the spread and orientation. The paper explains how scale and rotation produce this matrix.
The shape has no hard edge. Its contribution becomes smaller as we move away from the centre.
From photos to a scene
We first need to work out where the cameras were. Structure from Motion, using a tool such as COLMAP, matches features across overlapping photographs to estimate camera poses and a sparse set of 3D points.
The original method places initial Gaussians at those points, then repeats a fitting loop:
- Render the scene from one of the known camera positions.
- Compare the result with the real photograph.
- Use gradients to adjust the Gaussian parameters and reduce the error.
The optimization also changes the number of Gaussians. It clones or splits them where more detail is needed and removes ones that contribute too little.
This fitting happens separately for each captured scene. The photographs supervise the result; a detailed mesh is not required as a starting point.
Why “splatting”?
To draw a frame, the renderer projects the Gaussians onto the image plane, approximating each as a soft ellipse. These are the splats.
Several splats can cover the same pixel. We blend their colours in depth order, accounting for how much the nearer ones hide the ones behind them. The gsplat renderer documentation describes this process.
For two splats, with the first in front:
Here , , and are RGB colours. Each is the splat's effective opacity at this pixel, including its Gaussian falloff.
Suppose a red splat sits in front of a blue one, both with opacity 0.5. Red contributes 50%, blue contributes 25%, and the remaining 25% comes from the background. Swap their order and the colour changes.
The renderer groups splats into screen tiles, sorts them by depth, and blends them on the GPU. Once the scene has been fitted, we can move the camera and render new views without fitting it again.
What do we actually get?
We get a representation of how the captured scene looks from different directions. That makes it useful for exploring a place, but it does not give us a clean triangle mesh with surfaces ready for collision detection or new lighting.
Capture quality matters. COLMAP needs overlapping views with enough visual detail to match. Moving objects, reflections, and large blank surfaces can make that harder. The original 3DGS method also assumes a static scene, and poorly observed regions can produce stretched or floating artifacts.
Still, being able to take a collection of flat photographs and turn them into a scene with parallax is pretty cool. Move the camera sideways and nearby objects shift more than distant ones. The place starts to feel like a space again.