Your speakers send a tone. Your microphone picks up its reflections. Move your hand, and ETHER turns the Doppler shift into music.
- Play without touching. Hand motion shapes continuous pitch using your computer’s speakers and microphone.
- Find your sound. Adjust tone, glide and reverb, with C pentatonic, C major or free pitch.
- Choose your input. Select a microphone and see what calibration actually detects.
- Keep audio on your device. Microphone audio is processed locally, without recording or uploading.
The interface supports Chinese and English, with Chinese selected initially. Touch and keyboard controls are available as a fallback.
Try the live demo, or run locally with Node.js 22.13+ and npm. No account, API key or cloud configuration is needed locally.
git clone https://github.com/openaigames/ether-theremin.git
cd ether-theremin
npm ci
npm run devOpen the localhost URL printed in the terminal. Microphone capture requires HTTPS or localhost.
- Use built-in speakers and a physical microphone, without headphones.
- Start sonar, grant microphone access, and remain still until calibration finishes.
- Move your palm toward the speakers to raise pitch and away to lower it. Melody response holds the note for 2.2 seconds when motion stops, then fades gently. Its C4–C5 range and slower response make small movements easier to control.
- Adjust tone, glide, reverb and scale assist. Stop, Escape or moving the page into the background releases the microphone and stops the probe.
The microphone selector lets you choose a specific input. Automatic mode tries a built-in microphone if the default is a recognized virtual input. Browsers may hide device labels before permission is granted.
System volume and Probe level control the sensing tone; Output volume controls only the music. Calibration tests 18.5–21 kHz by default. The optional compatibility band tries 16–18 kHz only after the high band fails. These tones may be audible to people or pets. Start quietly and stop if uncomfortable.
Melody response + C pentatonic is the starting preset. It uses C, D, E, G and A, with note-boundary hysteresis to reduce jitter. Gesture response restores the faster C3–C6 range and quick fade; select Free pitch to play continuous glides without snapping.
- Click Listen before starting sonar to hear the original eight-note phrase. It plays at 80 BPM without microphone access; Stop and Escape cancel it.
- Start sonar and click Follow along. Practice selects Melody response and C pentatonic.
- Move gently toward the speakers to rise, away to fall. Hold each highlighted note for about half a second; practice follows your pace.
- Use End note to create a rest without recalibrating. The next detected motion sounds the instrument again. Stop releases the microphone.
The phrase is C4 E4 G4 A4 | G4 E4 D4 C4. A4 and the final C4 last two beats in the demo. The guide tracks stable notes, not rhythmic accuracy. Melody mode remains motion-based and does not measure hand distance.
For musical phrasing and traditional playing technique, see Carolina Eyck’s tutorials and Lydia Kavina’s lessons. Their antenna-based finger positions do not map directly to this sonar controller.
Open Sonar calibration diagnostics to see the actual input, sample rates, input level, reported audio processing and per-frequency stability.
| What you see | What to try |
|---|---|
| Input is almost silent | Select a physical microphone instead of a virtual input. Check system input volume and mute. |
| Input works, but the probe is weak | Check speaker routing, system volume and Probe level. Output volume changes only the music. |
| High-frequency calibration keeps failing | Try the optional compatibility band. On Mac, check microphone mode and disable Voice Isolation. |
| Device names are missing | Grant microphone permission first; the browser may hide labels until then. |
Browser-reported settings cannot rule out additional macOS or hardware processing. See Apple’s microphone mode guide.
Audio is processed only on the device: it is not recorded, uploaded or played back. Diagnostics and device identifiers stay in page memory. Only the theme preference is stored locally. Optional WebMCP tools expose state, input enumeration, configuration and stop actions; they never start audio or request microphone permission. state.message contains a stable translation key.
- Listen to the room. Measure a stationary background and find a usable carrier frequency.
- Listen for movement. Compare the frequency spread on either side of the carrier to estimate motion direction.
- Turn motion into pitch. Integrate the signal over time, then apply the selected tone, glide and reverb.
Calibration and signal processing
Inspired by Daniel Rapp’s doppler and SoundWave, CHI 2012. The engine uses linear spectral power, a stationary baseline, a second bandwidth scan for detached peaks, and time-based pitch integration. Calibration uses a multi-frame background median and requires four stable frames per accepted carrier. Both microphone and processing sample rates constrain the usable band.
This senses motion, not absolute distance, and does not track two hands independently. It borrows the theremin’s continuous pitch expression rather than its capacitive sensing. Results depend on hardware, browser and room reflections; synthetic tests cannot validate every acoustic environment.
npm run check # Tests, TypeScript and lint
npm run build # Build the Cloudflare Worker
npm run start # Preview the production build locallyContributions are welcome. See CONTRIBUTING.md for audio lifecycle, privacy and translation conventions.
Project map and i18n
app/: playing interface, styles and metadata.components/: diagnostics, instrument history and used UI primitives.lib/instrument.ts: audio, device selection and lifecycle.lib/signal.ts: signal processing and pitch integration.lib/melody.ts: stable scale snapping, sustain envelope and shared practice score.components/melody-coach.tsx: audition controls and guided note practice.lib/locales/zh.tsanden.ts: UI, metadata, statuses and errors.lib/instrument-errors.ts: browser failures mapped to stable message keys.lib/webmcp.ts: optional structured browser tools.tests/: synthetic signals and simulated audio devices.
Chinese defines the dictionary shape; English uses satisfies Locale to catch missing or misspelled keys. Add both translations for every new message. The engine emits typed MessageKey values, not display text. Keep translatable prose out of the engine and JSX. Brand names, source names, browser-reported device labels and tool protocol identifiers remain unchanged. Switching language does not recreate the audio engine.
Deployment and source export
The standalone source runs and builds without a Sites binding. For Sites hosting, configure your own .openai/hosting.json using .openai/hosting.example.json as a starting point. Never reuse another site’s project identifier.
npm run export:source -- ../ether-theremin-sourceThe allowlisted export excludes Git history, deployment bindings, dependencies, builds and temporary files. It checks for common credential patterns and local paths. The destination must not exist. Exporting does not create or publish a remote repository.
Thank you to the people who made this experiment possible:
- Emanuel Perez · @emanperez28 — the sonar scrolling demo on X that inspired ETHER’s acoustic gesture control.
- Thomas Kellogg · @oldnickels — the suggestion to make a theremin, which inspired ETHER’s musical direction.
- Daniel Rapp · doppler — the browser implementation of acoustic Doppler motion sensing, and the MIT-licensed bandwidth algorithm adapted in ETHER.
- SoundWave · CHI 2012 — the research foundation for sensing gestures with speakers and a microphone.
The instrument’s history section draws on the Bob Moog Foundation, Museums Victoria and the Clara Rockmore biography. Further references and asset credits are collected in Third-party notices.
Project-owned code is MIT licensed. Third-party materials retain their own licenses: doppler and shadcn (MIT), Manrope (OFL 1.1), and the museum photograph (CC BY 4.0). The main artwork was AI-generated for this project and is not historical imagery. See THIRD_PARTY_NOTICES.md for attribution, asset terms and historical references.
