Skip to content

Repository files navigation

chatsounds.metastruct.net

Get sounds into a chatsounds repo, entirely in the browser. Four tabs:

  • Extract, the helper: cut one long recording into named clips. It finds where each voice line starts and ends, transcribes it, names the clip after what is said, and gives you an editor to fix what the machine got wrong: trim, extend, split, merge, rename, add one it missed. Ends in a ZIP of .ogg files whose names are already legal chatsounds triggers.
  • Upload, the real pipeline: sort .ogg files into realms (the folder a sound belongs to in the repo). Realm names autocomplete from the live Metastruct repo, new ones are allowed, each realm is a drag-and-drop area, and every dropped file is checked against what the game can play (Vorbis, 44.1 kHz, mono or stereo) by reading its header. Filenames are folded to the trigger rules on the way in, since the filename is the trigger phrase and the repo's preprocessor rejects uppercase paths outright. Sign in with GitHub and the form opens the pull request itself: fork, branch, commit, PR, all from the browser.
  • Review, for the people who look after the repo: pick an open pull request, hear every sound in it, see what the checks flagged (too long, too much silence, wrong format), deny sounds with comments that land on the PR, approve or request changes on the whole thing, and merge. Restricted to accounts with push access; everyone else is told so.
  • Explore, the reference: every sound in the repo as one tree, realm by trigger by variation, with a play button on every row. It opens on the realm names alone, since that is a list a person can read; a realm unfolds when clicked, and searching unfolds whatever matches, name or trigger. Answers the question that comes before adding a sound, which is whether it is already there. Needs no sign-in. The copy button on a row yields a link that plays and names the sound wherever it is pasted.

Extract deliberately knows nothing about realms and Upload nothing about recordings: the bridge between them is a downloaded zip.

Built for neo-chatsounds and garrysmod-chatsounds.


Quick start

docker compose up -d --build

Then open http://localhost:8080.

The speech model downloads on first use (~80 MB for the default) and is cached by the browser afterwards.

Serve it over HTTPS. Browsers only grant SharedArrayBuffer, and with it multi-threaded WebAssembly, to cross-origin isolated pages, which requires a trustworthy origin. Over plain HTTP on a LAN address everything still works, just single-threaded and several times slower. localhost counts as trustworthy; 192.168.x.x does not.

Turning on GitHub sign-in

The Upload tab opens pull requests as the signed-in user. To enable it, register a GitHub OAuth app (any callback URL, it is never used), tick Enable Device Flow, and pass the app's client id:

GITHUB_CLIENT_ID=Ov23liAbCdEf... docker compose up -d --build

Left unset, the tab simply says sign-in is not available on this copy.

Why the device flow. A static page cannot hold the client secret the normal OAuth redirect flow needs, and self-hosters on localhost, LAN addresses and their own domains could never share one registered callback URL. The device flow needs neither: the page shows a short code, the user enters it at github.com/login/device, and the page polls for its token using only the public client id. The one thing github.com's OAuth endpoints lack is CORS headers, so nginx forwards /github/device/code and /github/oauth/token to github.com verbatim (see docker/nginx.conf.template); Vite's dev server does the same, with VITE_GITHUB_CLIENT_ID in .env.local supplying the id. The forwarder holds no secret and no state. Tokens are scoped to public_repo and live in the browser's localStorage until sign-out.

Sharing one sound. The server answers three routes besides the app itself.

  • /s/<realm>/<trigger>.ogg is the link the copy button yields: a small page carrying that one sound's Open Graph tags and a player, because a chat client's crawler reads tags out of HTML and never runs the app's JavaScript. It must not be served as audio/ogg just because its URL ends in .ogg, which is why its location empties nginx's type table.
  • /stream/<realm>/<trigger>.ogg proxies the file from raw.githubusercontent.com with content-disposition: attachment stripped, the header that makes a link to GitHub download a sound instead of playing it.
  • /mp4/<realm>/<trigger>.mp4 is the same sound as an MP4: the audio as mono 64k AAC, which is plenty for a preview of an already-lossy Vorbis file, under a black 400x144 frame. Discord draws its own controls over the bottom of the video, so a picture in the frame is only something nobody asked for sitting above a player, and 400x144 is the box Discord reserves for embedded media whatever the video says its size is: a flatter one buys no space back, it just leaves the difference empty underneath. Filling the box is what stops the embed resizing as it loads. Discord has never supported the OpenGraph audio tags and reserves its player embeds for a few whitelisted music services, so an MP4 behind og:video is the only way a link from anybody's domain plays inline. mp4d (docker/mp4d.mjs) builds one the first time it is asked for and nginx serves it off a volume every time after, so only sounds somebody actually shares are ever encoded.

The share page's only Open Graph tags are the player: a title, description and site name each render as another line above it, and that chrome is what makes an embed of a three second sound take over a channel. Its <title> stays for the browser tab, which is a different audience.

All three live in docker/nginx.conf.template and are mirrored in vite.config.ts for development. The path is reflected into the share page's markup and used as a filename, and nginx cannot escape anything, so the safety is the location's character class: every path in the repo is lowercase letters, digits, space and . _ - ! ( ) +, and a request carrying anything else matches no location and falls through to the app.

How the PR is made. Everything runs in the tab against the REST API: ensure a fork exists (creating one is idempotent and asynchronous, so it is polled), best-effort sync the fork with upstream, branch from its head, upload each sound as a base64 blob, build a tree and commit, then open the cross-repo pull request against master. Every step is idempotent or freshly named, so a failed run is safe to retry.

Reviewing

The Review tab gates on the repo itself: GET /repos/... reports the caller's permissions, and push (what a merge needs) is what seeing the page needs.

Opening a PR fetches its file list, then each sound's bytes from the head repo's blobs API (authenticated and CORS-clean, unlike raw URLs), and runs two kinds of checks. Format is the Upload rules again, identifyOgg on the header. The audio itself is decoded and audited (pipeline/audit.ts, pure and tested): over 30 s flags long, and more than 2 s of combined leading/trailing silence, or a sound more than a third under −45 dB RMS, flags silence. The checks are advisory; the play button outranks them.

Verdicts map onto what GitHub actually supports, one review at a time:

  • deny one sound asks for a comment; the first denial submits a changes-requested review naming the sound, later ones post as PR comments, so the author gets one red review and a thread rather than a stack of reds.
  • Deny all is one changes-requested review with one comment; Approve all is an approving review. GitHub refuses self-approval; the error is shown as-is.
  • Merge tries a merge commit and falls back to squash if the repo disallows it. The button disables itself when GitHub says the PR is not mergeable.

Files a PR touches outside sound/chatsounds/autoadd/ are listed separately as things this page cannot check, with a pointer to review them on GitHub.


Why the output looks the way it does

In chatsounds the filename is the trigger phrase, so the whole job is really "name these clips correctly". neo-chatsounds derives the trigger from the path:

key = chunk:lower():gsub("%.ogg$", ""):gsub("[%_%-]", " "):gsub("[%s\t\n\r]+", " "):Trim()

and the chat parser only strips " and ' from what a player types. A trigger containing any other punctuation is therefore unreachable: the filename keeps the character, the typed message loses it, and they never match. So the app strips punctuation up front, and everything it writes obeys:

Rule Why
.ogg, Vorbis the modern loader ignores every other extension, and GMod plays these through BASS, which needs a plugin for Opus
44.1 kHz, mono, -q:a 3 equivalent the spec in the addon's HOW TO ADD SOUNDS.txt
all-lowercase paths the legacy preprocessor rejects non-lowercase paths outright
[a-z0-9 ] triggers only anything else cannot be typed into chat and matched
variations numbered 01, 02, … the addon orders variations by URL, so 1, 10, 2 would shuffle :select(n)
never emits sh reserved by the addon for stopping playback

A name used once stays a flat file. The moment two clips share a name, both move into a folder and become numbered variations, which is exactly how the addon models "pick one of these", and is also the only way two clips can share a name at all:

hello there.ogg
get down/01.ogg
get down/02.ogg

That is the whole zip. No sounds/chatsounds/<pack>/ prefix, no index files, no repo_config.json snippet to copy by hand.

Why nothing above the clips. Every one of those needs to know which repository this is going to, and that is the publish step's business, not this screen's. The prefix, the repo_config.json line, and the two optional index files (list.msgpack and the legacy lua/chatsounds/lists_nosend/<realm>.lua, both keyed by realm) all came from a version of this page that tried to do the publishing on your behalf by telling you how. They have been removed rather than left switched off; a publish flow that has a GitHub token can regenerate any of them from the clip list, correctly, without asking.

The app has no concept of realms or speakers. It segments per voice line, and stops there.


How it works

file ─► decodeAudioData ──► OfflineAudioContext ─┬─► 16 kHz mono ──► silero VAD ──► speech intervals
        (main thread, native)                    │                └─► Whisper ──► word timings
                                                 └─► 44.1 kHz mono ──► every clip is cut from here
                                                        │
                    segmenter: VAD intervals ∩ word timings ──► voice lines
                                                        │
                    naming: transcript ──► trigger ──► collision-resolved paths
                                                        │
                                       libvorbis (WASM) ──► .ogg ──► zip

Neither detector is sufficient alone. Silero knows precisely where speech is but nothing about what it says; Whisper knows the words but its segment boundaries routinely glue two lines together or cut mid-word. So silero's intervals are the skeleton, Whisper's words hang off them, and the word timings are used only to decide where an over-long interval should be broken, at its most balanced internal pause. Boundaries are then snapped to the nearest local energy minimum, so cuts land in silence rather than clipping a syllable.

Because word timings are estimated while VAD boundaries are measured, a word may nudge a boundary by at most 250 ms. Without that cap one mistimed word stretches its line across the silence and swallows the next one, which is exactly the difference between hello / there i am a doctor and the correct hello there / i am a doctor.

A few consequences worth knowing:

  • Editing is instant. Clips are cut from the decoded master already in memory, so dragging a boundary never re-decodes anything. A clip is encoded lazily, on first play or export, and cached by a key derived from its bounds, gain and quality, so moving a boundary back reuses the previous render.
  • Playback is instant too. The editor plays time ranges of the master rather than a file per clip, which is also what makes previewing an extended clip possible: the audio outside the current bounds is already there.
  • Long files stay cheap to draw. The waveform comes from a precomputed 200 Hz envelope, so a 90-minute recording draws as fast as a 10-second one, and the zoomed clip editor slices that same envelope at 5 ms resolution.

Notable constraints

  • Everything is held in memory. A 90-minute recording is roughly 500 MB of decoded audio once resampled. Minutes-long voice-line dumps are the intended case; feature-length files will strain a tab.
  • A reload loses the work. There is nowhere to save it to.
  • mkv and avi cannot be decoded by any browser. The app says so and tells you the one-line ffmpeg remux to run. mp4/m4a/mov depend on AAC being available: Chrome and Edge ship it, Firefox borrows it from the system, and Chromium builds without proprietary codecs do not have it. When the decode fails the app hands you the ffmpeg line for that too.

The editor

The overview timeline is the main instrument, not a picture of one:

Gesture
click a clip select it
drag its edge trim or extend it
drag empty space add a clip
click elsewhere play from there

A clip drawn by hand has no words behind it, so it is snapped to the nearest quiet points and then transcribed on its own straight away. A clip with no name is the one thing here of no use at all, and the same one-clip transcription is what names the second half after a split.

Everything else the panel used to offer as a button is either a gesture now or a key. Two dozen nudge buttons said less than dragging the edge does:

space play the selected clip
j / k next / previous clip
[ ] move the start ±50 ms
{ } move the end ±50 ms
n add a clip at the playhead
s cut in two at the playhead
m join to the next clip
x delete
enter rename

What is left in the toolbar is adding a clip, searching, and the download. The bulk operations that were there (filter to the flagged ones, delete them all, match every clip's loudness) are gone: three buttons crowding out the two that matter. Clips the segmenter was unsure about still carry a tag, no words, long or short, since no words is how you spot one that still needs a name, and per-clip volume is still a slider.

The zoomed view beside the list stays, because on a ninety-minute recording a two-second clip is two pixels wide and there is nothing to grab on the overview. Same gestures, one clip at a time.


Models

Whisper runs through transformers.js. Four sizes are offered; base is the default.

Model Download Notes
tiny.en ~40 MB fastest, noticeably worse triggers
base ~80 MB the default
small ~250 MB better, wants WebGPU
large-v3-turbo ~800 MB best, effectively WebGPU-only

All four are the _timestamped exports specifically. Word-level timestamps come from Whisper's cross-attentions, and a model has to be exported with output_attentions=True for those to exist in the graph at all. The plain ONNX exports fail outright when asked for word timings, which would take the segmenter's basis for splitting long lines with them.

WebGPU is used when available and WebAssembly otherwise. The WASM path works but is several times slower; the upload screen says so when it detects no WebGPU, and names the setting to change for the browser you are actually using.

Firefox. WebGPU is on by default on Windows only. On Linux and macOS it is still behind dom.webgpu.enabled in about:config, and a restart. Nothing the page does can turn it on. about:support reports what the graphics stack settled on.

Which backend, and what happens when it fails

Runs on in the detection settings is automatic, WebGPU (force) or WebAssembly (CPU). Automatic takes the GPU wherever it can, except Firefox, where it grants an adapter and loads the model and then fails partway through inference inside onnxruntime's own buffer manager:

failed to call OrtRun(). ERROR_CODE: 1 … webgpu/buffer_manager.cc:553
Failed to download data from buffer: Mapping WebGPU buffer failed: Invalid buffer

Whisper is where that surfaces, being the only model here big enough to reach it. Forcing WebGPU still tries: the implementation moves quickly and this is a preference, not a lockout.

Any failure during transcription is then walked down a short ladder, each rung a real attempt on the same audio: GPU → CPU quantised → CPU full precision. The last is the combination nothing has been observed to reject, and also the largest and slowest, hence last.

Each rung needs a fresh worker, and that is not an optimisation. It is the mechanism. transformers.js funnels every session creation and every inference through a promise chain it never clears:

return apis.IS_WEB_ENV ? webInferenceChain = webInferenceChain.then(run) : run()

With no rejection handler, the first failure leaves that chain rejected forever, and every later call returns the same error without running anything. So a retry in the same worker is not slow, it is impossible, which is also why the app used to report the first dtype's error after appearing to try a second. The store therefore terminates the worker, rebuilds the 16 kHz working audio from the master (it was transferred, not copied), and starts over, saying so on the progress screen rather than silently rewinding the stage list.

Weights are fetched from the Hugging Face CDN and cached by the browser. That and the realm list are the only third-party requests the app makes. Everything else, including the fonts and the VAD model, is served from your own origin. To remove it entirely, mirror the model files and point env.remoteHost in src/pipeline/asr.ts at your copy.

onnxruntime and its WebAssembly

Both the VAD and Whisper run on onnxruntime-web, and src/pipeline/ort.ts exists to make sure that is one runtime with its binary on our own origin. Two things go wrong otherwise, and both surface as the same message, "no available backend found. ERR: [wasm] … failed to match magic number":

  • onnxruntime locates its .wasm by resolving the filename against the module that loaded it, which under a bundler is a path nothing serves. The dev server answers it with index.html, which then fails to compile as WebAssembly.
  • transformers.js, finding the path unset, points it at jsdelivr, turning a binary we already ship into a 23 MB third-party download.

So ort.ts names both files explicitly, as URLs Vite emits as assets, and picks the pair that matches the backend in use: the asyncify build (23 MB) can drive WebGPU, while the plain build (13 MB) is CPU-only and is what a browser without a GPU is given. onnxruntime-web is also pinned in package.json to the exact build transformers.js depends on, and forced on it via overrides, because two copies means two env objects: configuring one leaves the other on the CDN. Bump that pin whenever transformers.js is bumped.


Development

cd frontend
npm install
npm run dev     # http://localhost:5173
npm test        # 125 unit tests
npm run build

The tests cover the parts that have to be exactly right: the trigger rules (including a round-trip through a reimplementation of the addon's own key derivation, so the names written to disk survive what the loader does to them), the segmenter's boundary and splitting logic, the envelope reader, the zip layout, where a hand-drawn clip is allowed to land, the ogg header parser the Upload form screens files with (verified against real ffmpeg output as well as crafted bytes), realm-name folding, the silence and length audit the Review tab runs, that onnxruntime is told where both halves of its WebAssembly are, and that the backend fallback ladder always terminates.

Layout

frontend/src/
├── pipeline/     decode · vad · asr · segmenter · naming · encode · pack
│                 plus ogg (what is really inside an .ogg, from its header),
│                 gpu (is there a WebGPU adapter, and why not), ort (where
│                 onnxruntime finds its own WebAssembly) and attempts (what to
│                 try next when the backend fails)
├── workers/      the VAD/ASR/encode worker -- Web Audio is main-thread only,
│                 so decoding stays outside it and everything else moves in
├── store/        useJob (Extract) · useUpload · useGithub · useReview ·
│                 useExplore · worker client
├── lib/          github (device flow · fork · PR) · realm list · sound index
│                 and its tree · gaps · icons
├── components/
│   ├── extract/  start · processing · editor · waveform · clip editor
│   ├── upload/   the realm form: combobox · drop area · sign-in and PR
│   ├── review/   PR picker · sound tree · verdicts and merge
│   ├── explore/  the whole repo as a windowed tree, filter and share
│   └──           Navbar (the four tabs) · Icon · GithubSignIn
└── styles/       design tokens taken from metastruct.net

Design

The interface follows metastruct.net: its navbar and logo, a #212121 page with #171717 chrome and #4a4a4a panels floating on a single soft dark halo, Open Sans at a 14px root, and a two-accent system that is the load-bearing idea: teal #09b387 means you can act on this, purple #7d3b80 means you are interacting with this. In the segment list that distinction does real work. Open Sans is self-hosted and the icons are inlined MDI paths, so the page itself pulls nothing from a CDN.

Three control heights, and no fourth

Every button, input, select and tag takes its height from --control-h (2.5rem), --control-h-sm (2rem) or --control-h-xs (1.5rem, tags only). In rem, never em.

That last part is the whole point. These heights used to be written in em, which makes a height a function of the element's own font size, so every variant that set smaller type silently got a smaller box: .button.is-small at 0.85rem came out 23.8px rather than the 28px it read as, an icon button whose glyph was set to 0.75rem came out 26px, and the download button beside it came out 24px. Seven controls, six heights, none of them chosen. A variant may change the type size; the box comes from the token.

The scale means: 2.5rem for a control standing on its own or in a form, 2rem for dense rows and toolbars, where every control in one row shares a height, 1.5rem for tags. An is-icon button is a square of its own height whatever the glyph inside is sized at.

About

Wrapper and helpers service to make it easier to manage Metastruct chatsounds.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Packages

Contributors

Languages