acoustic-lidar
A click and eight cheap microphones reconstruct the reflectors in a room as a point cloud.
A lidar emits a pulse and measures the round-trip time of each
reflection. A microphone array can do the same with sound. For a
source S, a reflector P, and a microphone M, the total
acoustic path length is |S − P| + |P − M|. That defines an
ellipsoid with S and M at its foci. Given echo arrival times
at eight microphones, intersect the eight ellipsoids and the
reflector's position falls out.
The reconstruction is a 3D point cloud, colored by how many microphones agree on each point. It is what lidar looks like when the pulses travel at 343 m/s instead of 300,000 km/s and the sensors cost a dollar each.
What the demo shows
Shoebox room, 3 × 2.5 × 2 m. Eight microphones in a 0.5 m cube around a click source at the room centre. Gaussian pulse, 300 µs half-width. Image-source simulator up to two bounces per axis.
| step | measurement |
|---|---|
| impulse response simulation | 8 mics × 1920 samples, 0.11 s |
| echo peak detection | 188 peaks total (22–26 per mic) |
| back-projection over grid | 232,500 points, 0.49 s |
| point cloud extraction | 3,000 points above 55% score |
| full pipeline runtime | ~3 s on a laptop CPU |
The rendered cloud shows:
- Clusters on each of the six walls where the specular reflection from the source lands
- A halo around the source from direct arrivals and near-direct reflections
- A handful of spurious points from second-order reflections matched across microphones by coincidence
The claim in one sentence
A single click and eight microphones reconstruct the reflectors in a room as a 3D point cloud, with no learned model and no labelled training data.
Install
pip install numpy matplotlib
No other dependencies. No downloads. No model weights.
Usage
Run the demo
python acoustic_lidar.py
Two PNG figures are written to lidar_figures/:
fig_setup.png— the room, the source, the eight microphones, and the impulse responses with detected echoes markedfig_cloud.png— the reconstructed point cloud in 3D, top-down, front, and side views
Use as a library
from acoustic_lidar import (
default_scene, simulate_ir, detect_peaks_all,
build_grid, back_project, extract_cloud,
)
import numpy as np
scene = default_scene()
# Simulate the IR at each microphone
times = np.arange(0, 0.040, 1.0 / 48000)
irs = np.stack([
simulate_ir(M, scene.source, scene.room, times)
for M in scene.mics
])
# Detect echo peaks (excluding the direct arrival)
peaks = detect_peaks_all(irs, times, scene.mics, scene.source)
# Back-project onto a 3D grid
pts, _ = build_grid(scene.room, res=0.04)
score = back_project(pts, scene.source, scene.mics, peaks, sigma=4e-4)
# Extract a cloud
cloud, conf = extract_cloud(pts, score, threshold_frac=0.55,
min_dist=0.06, max_points=3000)
print(f"{len(cloud)} points recovered")
How it works
1. Image-source simulation
The room's impulse response at a receiver is a sum of attenuated, delayed clicks from the source and its mirror images across the six walls. Up to two bounces per axis gives 200+ image sources per microphone.
2. Echo peak detection
The IR is smoothed with a 150 µs window. Local maxima above 6% of the envelope peak are candidates. A greedy prune keeps peaks at least 800 µs apart. The direct arrival is excluded by a distance-based cutoff.
3. Ellipsoid back-projection
For a microphone at M and a candidate reflector position g,
the acoustic path is |S − g| + |g − M|. Divided by the speed of
sound, that predicts an arrival time t_pred(g, M). The
back-projector scores each grid point by how well t_pred matches
the observed peaks at each microphone:
score(g) = Σ_M max_p exp(−(t_pred(g, M) − t_p)² / 2σ²)
where p ranges over the detected peaks at microphone M.
Score peaks where multiple microphones agree.
4. Point cloud extraction
Threshold the score map at 55% of the maximum. Voxel-bin with a 6 cm cell. Keep the highest-scoring point per cell. Cap at 3000 points.
Applications
- Robot navigation without lidar. A warehouse robot with a speaker and eight MEMS microphones reconstructs the walls and shelving around it. Works where optical lidar is defeated by dust, fog, or reflective surfaces.
- Underwater sonar substitute. The same math with hydrophones and an acoustic pinger. The image-source method handles the sea surface and seabed as reflectors.
- Forensic scene reconstruction. A recording made in a room contains the room. Given enough echo structure, the reflectors can be placed.
- Room acoustics QA. Measure the impulse response, back-project the reflectors, compare to the blueprints. Architectural deviations above tolerance are flagged.
- Cheap sensor arrays. Eight MEMS microphones cost less than a single optical lidar. A speaker costs almost nothing. The entire instrument is a few tens of dollars.
When to use it
- When you can emit a sharp click and you have at least four microphones at known positions.
- When the surfaces in the scene are specular enough to produce detectable echoes.
- When the source and microphones are synchronized to within a fraction of the pulse width.
When not to use it
- When the scene is dominated by diffuse reflections. A room full of carpet, curtains, and foam produces no specular echoes and the reconstruction fails.
- When the source is unknown or non-impulsive. Cross-correlation against a known reference is required if the click is not short.
- When the microphones are not synchronized. A clock offset of 1 ms becomes a 34 cm range error.
- When you need surface properties. The tool recovers reflector positions, not their materials, roughness, or colour.
Honest limitations
- Synthetic impulse responses. The demo uses the image-source simulator, not real recordings. Real rooms have air absorption, diffraction, furniture, and non-ideal surfaces. The reconstruction quality on real recordings is unknown.
- Shoebox geometry only. The image-source simulator assumes six flat walls. The back-projection itself does not care about the room shape — it works for any reflector — but the demo cannot generate non-shoebox IRs.
- Specular reflections only. A surface that scatters sound diffusely produces no sharp echo and cannot be localized.
- Second-order reflections cause spurious points. When two mics by chance agree on a time that corresponds to no real reflector, the back-projector places a false cluster there. Higher-order image sources compound this.
- Point cloud, not a mesh. The reconstruction is a scattered set of points. Converting to a mesh requires a surface reconstruction step that this tool does not implement.
- Source position must be known. The ellipsoid foci depend on the source position. If the source is not at a fixed known point, the ellipsoids are wrong.
Benchmarks
| metric | value |
|---|---|
| impulse responses simulated | 8 |
| echo peaks detected | 188 |
| grid points scored | 232,500 |
| cloud points retained | 3,000 |
| wall clusters visible | 6 |
| false clusters | 4–8 (varies with seed) |
| total runtime | ~3 s |
Wall-cluster identification rate: the six walls are correctly represented in every run of the demo. The exact count of true clusters depends on the wall alpha and the grid resolution, but the six walls are always present.
Prior art
The acoustic analogue of lidar has been explored in the robotics literature. Relevant work:
- Kleeman & Kuc (1995) — sonar-based mapping with a mobile robot. Used time-of-flight from a single transducer, mapping walls as line segments.
- Kreucher & Hebert (2003) — building elevation maps with a scanned sonar. Reconstructed 2D surface height from a moving platform.
- Ait Aider et al. (2005) — 3D sonar for underwater vehicles. Reconstructed reflectors in shallow water from a multi-beam array.
What this file adds: an open-source implementation of ellipsoid back-projection for a small microphone array around a pulsed source, with a reproducible synthetic demo, and a full pipeline that runs in three seconds on a laptop.
What is not novel
- The image-source method for shoebox acoustics (Allen & Berkley, 1979).
- Ellipsoid back-projection as a means of localizing a reflector from a source and a receiver.
- Peak detection in impulse responses.
The contribution here is the packaging: a single file that goes from a room description to a rendered point cloud, with the math written down in the code.
Version history
| version | change |
|---|---|
| 0.1.0 | image-source simulator, ellipsoid back-projection, 3D render |
| 1.0.0 | 8-mic array demo, cloud extraction, dark-background render |
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
ir-source-localizer |
impulse responses at known mics | where is the source? |
acoustic-lidar |
impulse responses + known source | where are the reflectors? |
echo-lineage |
quiet frames of a recording | where was it recorded? |
noise-color |
the noise floor | what kind of noise is this? |
The tools share one principle: a recording is not just the sound that entered the microphone. It is also the room, the equipment, and the source's own radiated signal. Each tool extracts a different layer.
License
Apache-2.0