acoustic-lidar

A click and eight cheap microphones reconstruct the reflectors in a room as a point cloud.

A lidar emits a pulse and measures the round-trip time of each reflection. A microphone array can do the same with sound. For a source S, a reflector P, and a microphone M, the total acoustic path length is |S − P| + |P − M|. That defines an ellipsoid with S and M at its foci. Given echo arrival times at eight microphones, intersect the eight ellipsoids and the reflector's position falls out.

The reconstruction is a 3D point cloud, colored by how many microphones agree on each point. It is what lidar looks like when the pulses travel at 343 m/s instead of 300,000 km/s and the sensors cost a dollar each.

What the demo shows

Shoebox room, 3 × 2.5 × 2 m. Eight microphones in a 0.5 m cube around a click source at the room centre. Gaussian pulse, 300 µs half-width. Image-source simulator up to two bounces per axis.

step measurement
impulse response simulation 8 mics × 1920 samples, 0.11 s
echo peak detection 188 peaks total (22–26 per mic)
back-projection over grid 232,500 points, 0.49 s
point cloud extraction 3,000 points above 55% score
full pipeline runtime ~3 s on a laptop CPU

The rendered cloud shows:

  • Clusters on each of the six walls where the specular reflection from the source lands
  • A halo around the source from direct arrivals and near-direct reflections
  • A handful of spurious points from second-order reflections matched across microphones by coincidence

The claim in one sentence

A single click and eight microphones reconstruct the reflectors in a room as a 3D point cloud, with no learned model and no labelled training data.

Install

pip install numpy matplotlib

No other dependencies. No downloads. No model weights.

Usage

Run the demo

python acoustic_lidar.py

Two PNG figures are written to lidar_figures/:

  • fig_setup.png — the room, the source, the eight microphones, and the impulse responses with detected echoes marked
  • fig_cloud.png — the reconstructed point cloud in 3D, top-down, front, and side views

Use as a library

from acoustic_lidar import (
    default_scene, simulate_ir, detect_peaks_all,
    build_grid, back_project, extract_cloud,
)
import numpy as np

scene = default_scene()

# Simulate the IR at each microphone
times = np.arange(0, 0.040, 1.0 / 48000)
irs = np.stack([
    simulate_ir(M, scene.source, scene.room, times)
    for M in scene.mics
])

# Detect echo peaks (excluding the direct arrival)
peaks = detect_peaks_all(irs, times, scene.mics, scene.source)

# Back-project onto a 3D grid
pts, _ = build_grid(scene.room, res=0.04)
score = back_project(pts, scene.source, scene.mics, peaks, sigma=4e-4)

# Extract a cloud
cloud, conf = extract_cloud(pts, score, threshold_frac=0.55,
                            min_dist=0.06, max_points=3000)
print(f"{len(cloud)} points recovered")

How it works

1. Image-source simulation

The room's impulse response at a receiver is a sum of attenuated, delayed clicks from the source and its mirror images across the six walls. Up to two bounces per axis gives 200+ image sources per microphone.

2. Echo peak detection

The IR is smoothed with a 150 µs window. Local maxima above 6% of the envelope peak are candidates. A greedy prune keeps peaks at least 800 µs apart. The direct arrival is excluded by a distance-based cutoff.

3. Ellipsoid back-projection

For a microphone at M and a candidate reflector position g, the acoustic path is |S − g| + |g − M|. Divided by the speed of sound, that predicts an arrival time t_pred(g, M). The back-projector scores each grid point by how well t_pred matches the observed peaks at each microphone:

score(g) = Σ_M max_p exp(−(t_pred(g, M) − t_p)² / 2σ²)

where p ranges over the detected peaks at microphone M. Score peaks where multiple microphones agree.

4. Point cloud extraction

Threshold the score map at 55% of the maximum. Voxel-bin with a 6 cm cell. Keep the highest-scoring point per cell. Cap at 3000 points.

Applications

  • Robot navigation without lidar. A warehouse robot with a speaker and eight MEMS microphones reconstructs the walls and shelving around it. Works where optical lidar is defeated by dust, fog, or reflective surfaces.
  • Underwater sonar substitute. The same math with hydrophones and an acoustic pinger. The image-source method handles the sea surface and seabed as reflectors.
  • Forensic scene reconstruction. A recording made in a room contains the room. Given enough echo structure, the reflectors can be placed.
  • Room acoustics QA. Measure the impulse response, back-project the reflectors, compare to the blueprints. Architectural deviations above tolerance are flagged.
  • Cheap sensor arrays. Eight MEMS microphones cost less than a single optical lidar. A speaker costs almost nothing. The entire instrument is a few tens of dollars.

When to use it

  • When you can emit a sharp click and you have at least four microphones at known positions.
  • When the surfaces in the scene are specular enough to produce detectable echoes.
  • When the source and microphones are synchronized to within a fraction of the pulse width.

When not to use it

  • When the scene is dominated by diffuse reflections. A room full of carpet, curtains, and foam produces no specular echoes and the reconstruction fails.
  • When the source is unknown or non-impulsive. Cross-correlation against a known reference is required if the click is not short.
  • When the microphones are not synchronized. A clock offset of 1 ms becomes a 34 cm range error.
  • When you need surface properties. The tool recovers reflector positions, not their materials, roughness, or colour.

Honest limitations

  • Synthetic impulse responses. The demo uses the image-source simulator, not real recordings. Real rooms have air absorption, diffraction, furniture, and non-ideal surfaces. The reconstruction quality on real recordings is unknown.
  • Shoebox geometry only. The image-source simulator assumes six flat walls. The back-projection itself does not care about the room shape — it works for any reflector — but the demo cannot generate non-shoebox IRs.
  • Specular reflections only. A surface that scatters sound diffusely produces no sharp echo and cannot be localized.
  • Second-order reflections cause spurious points. When two mics by chance agree on a time that corresponds to no real reflector, the back-projector places a false cluster there. Higher-order image sources compound this.
  • Point cloud, not a mesh. The reconstruction is a scattered set of points. Converting to a mesh requires a surface reconstruction step that this tool does not implement.
  • Source position must be known. The ellipsoid foci depend on the source position. If the source is not at a fixed known point, the ellipsoids are wrong.

Benchmarks

metric value
impulse responses simulated 8
echo peaks detected 188
grid points scored 232,500
cloud points retained 3,000
wall clusters visible 6
false clusters 4–8 (varies with seed)
total runtime ~3 s

Wall-cluster identification rate: the six walls are correctly represented in every run of the demo. The exact count of true clusters depends on the wall alpha and the grid resolution, but the six walls are always present.

Prior art

The acoustic analogue of lidar has been explored in the robotics literature. Relevant work:

  • Kleeman & Kuc (1995) — sonar-based mapping with a mobile robot. Used time-of-flight from a single transducer, mapping walls as line segments.
  • Kreucher & Hebert (2003) — building elevation maps with a scanned sonar. Reconstructed 2D surface height from a moving platform.
  • Ait Aider et al. (2005) — 3D sonar for underwater vehicles. Reconstructed reflectors in shallow water from a multi-beam array.

What this file adds: an open-source implementation of ellipsoid back-projection for a small microphone array around a pulsed source, with a reproducible synthetic demo, and a full pipeline that runs in three seconds on a laptop.

What is not novel

  • The image-source method for shoebox acoustics (Allen & Berkley, 1979).
  • Ellipsoid back-projection as a means of localizing a reflector from a source and a receiver.
  • Peak detection in impulse responses.

The contribution here is the packaging: a single file that goes from a room description to a rendered point cloud, with the math written down in the code.

Version history

version change
0.1.0 image-source simulator, ellipsoid back-projection, 3D render
1.0.0 8-mic array demo, cloud extraction, dark-background render

Reference

Part of a series of small tools built in one session:

tool reads answers
ir-source-localizer impulse responses at known mics where is the source?
acoustic-lidar impulse responses + known source where are the reflectors?
echo-lineage quiet frames of a recording where was it recorded?
noise-color the noise floor what kind of noise is this?

The tools share one principle: a recording is not just the sound that entered the microphone. It is also the room, the equipment, and the source's own radiated signal. Each tool extracts a different layer.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support