Skip to content
View djzoom's full-sized avatar

Organizations

@HYDAOCAST

Block or report djzoom

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
djzoom/README.md

加菲众 Garfield Zhong Wang

Radio DJ · Speech & Audio ML · macOS Engineering · Founder & CEO

I build audio and speech AI from first principles — and ship what I learn.

Website · Email · X


About

I work on audio and speech AI from first principles. My main project is AURORA, a from-scratch course that rebuilds the audio-AI stack by hand in NumPy — from a single sine wave to Whisper (FFT, MFCC, backprop, attention, CTC, RAG), no black-box libraries. It's open-core: Lesson 1 is free and open source. I designed and directed the curriculum; it's built with AI-assisted pair programming. And I ship what I learn: TalkTalk, a macOS teleprompter built in Swift, is live (StoreKit/IAP, code signing, notarization).

Before I wrote software, I spent nearly two decades as a broadcaster designing audio for millions of listeners. That gives me an unusual edge in speech ML: I know what "good" sounds like — timing, prosody, pacing — before the metrics do.

Current Focus

  • AURORA — a from-scratch audio-AI course: 99 lessons, every algorithm hand-written in NumPy (FFT · STFT · Mel · MFCC · attention · CTC · RAG), validated against reference implementations. Open core — Lesson 1 free.
  • TalkTalk — shipped macOS teleprompter (Swift), 300+ commits, two major architecture rewrites. Exploring breath-group-based line breaking and broadcast-grade prosody metrics.
  • LiveCaption — exploring real-time, low-latency bilingual (zh/en) speech interfaces.

Selected Projects

  • MU5735 — reconstructing a tragedy through open aviation data and 3D visualization.
  • kimchi-fermentation-simulator — food science as an interactive numerical model.
  • polymarket-trends — live dashboards tracking prediction-market movement.
  • 0xgarfield-home — personal site and publishing hub.

Research Interests

  • Streaming / on-device ASR
  • Prosody and paralinguistics
  • Bilingual (zh/en) speech systems
  • Human-in-the-loop audio tooling

Background

Nearly two decades in professional broadcasting (HIT FM / China Radio International, Phoenix URadio). Judge, 47th International Emmy Awards (Long-Form Documentary Sound). Chemistry degree — the quantitative habits stuck.

Links

Popular repositories Loading

  1. MU5735 MU5735 Public

    MU5735 3D Flight Reconstruction — ADS-B + FDR data visualization (Three.js)

    HTML 109 20

  2. kimchi-fermentation-simulator kimchi-fermentation-simulator Public

    Interactive Korean kimchi fermentation simulator with scientific models (Gompertz/Arrhenius). Trilingual UI. 김치 발효 시뮬레이터 | 辣白菜发酵模拟器

    JavaScript 5

  3. AURORA AURORA Public

    From-scratch audio-AI research system + 99-lesson bilingual course — DSP · ASR · Music · LLM/RAG, every core hand-written, no API wrappers.

    Python 4

  4. polymarket-trends polymarket-trends Public

    A real-time interactive dashboard that visualizes trending prediction markets from Polymarket. Tracks daily changes in probability, volume, and sentiment, generating an intuitive "market heat index…

    JavaScript 3

  5. starry-night-flow starry-night-flow Public

    HTML 2

  6. EB1A EB1A Public

    Python 1 1