ChordMini ChordNet (2E1D) β classifier + CQT plan (ONNX / WebGPU)
ONNX export of the ChordMini chord recognizer (ChordNet "2E1D", 170-class
large vocabulary), packaged for the musetric packages/ai runtime
(onnxruntime-web on WebGPU).
The graph is the classifier only: it takes log-CQT feature windows and
returns per-frame chord logits. Feature extraction is deliberately not baked
in β the host computes a recursive constant-Q transform on WebGPU and hands the
result over as a GPU buffer, so no features cross back to the CPU. This is not a
drop-in audio -> chords model.
mono PCM @ 22050 Hz (arithmetic-mean downmix β see Limitations)
-> WebGPU recursive CQT -> log(|CQT| + 1e-6) features [T, 144]
-> pad/window, groups of 16 -> [16, 108, 144] per run
-> chordnet.onnx -> logits [16, 108, 170] per run
-> WebGPU smoothing + argmax -> chord indices [T]
cqt-plan.bin ships with the model because it defines the features the graph
expects: the octave schedule, the sparse per-octave FFT basis and the resampling
FIR, baked from librosa 0.11.0. Model and plan are a matched pair β a release
therefore carries a hashable feature-extraction contract instead of an implicit
one.
Normalization ((x - mean) / (std + 1e-8)) is inside the graph. CQT, windowing,
smoothing and argmax stay in the host so their GPU buffers stay reusable.
Intended uses & limitations
Intended:
- Chord recognition over music, as a stage in an audio pipeline.
- Client/edge inference via WebGPU through
onnxruntime-web.
Out of scope:
- Standalone use without a host that computes librosa-equivalent log-CQT
features, windows them to 108 frames, and applies smoothing + argmax to the
logits (see
musetricpackages/aiandpackages/cqt). - Use in other training frameworks β this is an inference-only export.
Limitations:
- The graph is static at 16 windows of 108 frames and 144 bins. Feed a
track in groups of 16 windows, pad the last group with zero windows and trim
its logits back. Windows never interact inside the model, so padding changes
no real logit. Any other batch fails inside the first rewritten
Transpose. - The features must be librosa-equivalent. Substituting a different CQT is
not free: an nnAudio
CQT1992v2stand-in correlates at ~0.998 yet still costs ~1.2% of frames end to end. Use the shipped plan. - The model is gain-sensitive. It was trained on
librosa.load's arithmetic mean downmix(L+R)/2.ffmpeg -ac 1uses an energy-preserving rematrix(L+R)/sqrt(2), i.e. a factor of β2, whichlog(|CQT| + 1e-6)turns into a constantlog(β2) = 0.347offset on every feature β afterstd = 1.719a uniform+0.20shift, enough to flip frames near a decision boundary. Downmix as the arithmetic mean. - Its
idx_to_chordcheckpoint map differs from the reference runner'sidx2voca_chord()on 70 of 170 indices, in enharmonic spelling only (Db:minvsC#:min).config.jsonships the runner's vocabulary. - Training-data provenance of the upstream checkpoint is not documented here.
How to use
The session runs the classifier; the host supplies features and consumes
logits.
import * as ort from 'onnxruntime-web/webgpu';
import { createCqt, verifyCqtPlanArtifact } from '@musetric/cqt/gpu';
const session = await ort.InferenceSession.create('chordnet.onnx', {
executionProviders: ['webgpu'],
preferredOutputLocation: { logits: 'gpu-buffer' },
});
const device = await ort.env.webgpu.device;
// cqt-plan.bin; verifies the payload against the SHA-256 it carries.
const plan = await verifyCqtPlanArtifact(new Uint8Array(planBytes));
const cqt = createCqt(device).get({ input: pcm, output: features, sampleCount, plan });
// cqt.run(encoder) writes log features [T, 144]; pad T up to a multiple of 108.
// Copy each group of 16 windows into `groupFeatures` and run it.
const input = ort.Tensor.fromGpuBuffer(groupFeatures, {
dataType: 'float32',
dims: [16, 108, 144],
});
const { logits } = await session.run({ features: input });
// logits: float32 [16, 108, 170] per group; concatenate the groups, drop the
// padding windows -> uniform 9-frame smoothing -> argmax -> indices
See the musetric packages/ai host code for the full CQT, smoothing/argmax and
segment-grouping pipeline.
Known defect: onnxruntime WebGPU on Adreno 6xx, and the static rewrite (2026-09-11)
Strict on-device verification (one recorded input, wasm EP of the same build as
the bit-exact reference, comparator validated by negative controls) showed the
WebGPU EP returns silently wrong logits for the original graph on Adreno
660 (OnePlus 9RT): 95.7% of chord logits mismatched the reference
(maxAbs 0.58, nondeterministic across runs of one session). The mechanism is
the shared-tile Transpose kernel of ORT's WebGPU EP: this graph contains 111
Transpose nodes, most of them (1,0,2) over tensors whose leading window
axis collapses to rank-2-swapped - exactly the broken kernel family (one-line
ORT fix pending upstream, transpose.cc tile + 1 -> + 8). Desktop NVIDIA
is bit-exact on the same input and build.
chordnet.onnx now ships the static rewrite: all 79 affected Transposes
become Reshape -> Gather(const int32 idx) -> Reshape (exact permutations of
values), everything else is untouched. Verified on 9RT WebGPU against the
wasm EP of the same file: 0 mismatched logits (maxAbs 1e-5), deterministic,
in both storageBufferCacheMode: 'simple' and the default; faithful to the
original graph within 4e-6 (verified in ORT CPU). Full evidence: the musetric
plan gpu/adreno660-strict-2026-09-11.md.
That rewrite is static: every Gather index and Reshape target is baked for the
captured batch. The first revision was built for one window but still declared
a dynamic W, so any host feeding a whole track failed and the only working
feed was one window per run, where the fixed cost of each WebGPU run dominates.
This revision is built for 16 windows and declares exactly
[16, 108, 144]. 16 is below the cap that matters: the frequency encoder runs
Softmax over [W * 108, 8, 2, 2], and onnxruntime's WebGPU Softmax writes
past its output once a dispatch has more than 65535 rows, which a batch of 38
windows or more would reach.
Checked against the wasm EP of the same file, one input of 26 and 60 windows:
| runtime | max |Ξlogit| | chord argmax | 26 windows, 1 β 16 per run | 60 windows |
|---|---|---|---|---|
| OnePlus 9RT, Adreno 660, both cache modes | 5.7e-6 | identical | 0.58 s β 0.53 s | 1.32 s β 1.06 s |
| desktop NVIDIA, D3D12 | 6.7e-6 | identical | 0.31 s β 0.03 s | 0.61 s β 0.06 s |
| Galaxy S24 Ultra, Adreno 750 | 1.99 | 1 of 2808 frames differs | 0.30 s β 0.18 s | 1.2 s β 0.35 s |
In ORT CPU this revision matches the original graph within 7e-6 at 1, 5, 16, 20 and 60 windows fed in padded groups, with identical argmax.
Known defect: onnxruntime WebGPU MatMul on Adreno 750 (2026-09-20)
The Galaxy S24 Ultra row above was the second defect. There the WebGPU EP
computes a float32 MatMul wrongly for two shapes this graph uses, by units
and deterministically, while wasm on the same device is right: a
four-dimensional product with the extents of the chord attention, and a product
whose output width is not a multiple of four, which the classifier
[16, 108, 144] x [144, 170] has. Measured on that device against wasm, a
single node of the classifier shape differs by 4.2 at 170 columns and by 1.9e-6
at 168 or 172.
chordnet.onnx now folds the two batch axes of every four-dimensional MatMul
into one and pads the classifier weight to 172 columns, slicing the result back.
Both rewrites are exact, and they are what makes the graph agree with wasm on
that GPU. Checked against the wasm EP of the same file, 16 windows of
deterministic features:
| runtime | previous revision | this revision |
|---|---|---|
| Galaxy S24 Ultra, Adreno 750, Chrome 153 | max |Ξ| 3.13, 386 of 1728 frames choose another chord | max |Ξ| 5.7e-6, no frame changed |
| OnePlus 9RT, Adreno 660, Chrome 152 | 5.3e-6, 753 ms | 5.7e-6, 635 ms |
| desktop NVIDIA, Chrome 153 | 5.7e-6, 933 ms | 6.7e-6, 774 ms |
In ORT CPU this revision matches the exported graph within 5.7e-6, the same as the previous one, with identical argmax.
The time attention on the vec4 kernel (2026-09-30)
The cause of the Adreno 750 defect is now known (musetric#891): the
driver miscompiles the scalar MatMul kernel that onnxruntime's WebGPU EP
generates for two- and four-dimensional products, while its three-dimensional
variant comes out right, which is why the fold above works. The EP takes the
scalar kernel whenever K or N is not a multiple of four, as in the time
attention of this graph, [16, 8, 108, 18] x [16, 8, 18, 108] and
[16, 8, 108, 108] x [16, 8, 108, 18].
This revision takes the time attention off that kernel: K and N are padded with zeros to 20, which is the vec4 kernel, and the extra output columns are sliced off. The two-row frequency attention keeps the fold, because onnxruntime runs products of eight rows or fewer on a vec4 kernel that is wrong on Adreno 660. Both are exact.
Checked against ORT CPU on the first 16 windows of the CQT features of the instrumental stem of Rxbyn, "Bad Side" (https://www.jamendo.com/track/1556580/bad-side, CC BY 3.0, from JamendoLyrics), as the musetric parity run computes them; onnxruntime-web 1.30.0 with musetric's WebGPU provider and session options, median of 5 runs with the logits read back:
| runtime | previous revision | this revision |
|---|---|---|
| Galaxy S24 Ultra, Adreno 750, Chrome 153 | max |Ξ| 4.8e-6, no frame changed, 101.2 ms | 4.8e-6, no frame changed, 97.8 ms |
| OnePlus 9RT, Adreno 660, Chrome 153 | 4.8e-6, no frame changed, 286.5 ms | 4.8e-6, no frame changed, 283.4 ms |
| desktop NVIDIA RTX 3060 Laptop, Chrome 153 | 4.8e-6, no frame changed, 24.0 ms | 4.8e-6, no frame changed, 20.7 ms |
Without any MatMul rewrite the same graph differs there by 2.07 on the
Adreno 750, with 94 of 1728 frames choosing another chord.
Rebuild recipe: capture_shapes.py --ops Transpose,MatMul --windows 16 --shape 16,108,144 --input features + rewrite_static_adreno.py --vec4-attention, then
dispatch_rows_audit.py, in musetric-toolkit scripts/onnx/.
Files
| File | Size | SHA256 |
|---|---|---|
chordnet.onnx |
17,080,550 B | 1faee8e1cf168300afe0f47517e76836ac200c3841b0fb430fad7dfea5ad6394 |
config.json |
3,009 B | 1f26c11ebea51ec08f12e813eb213a729fa0ecc407ac7632dfdc7bad67e65aa4 |
cqt-plan.bin |
23,896 B | c31f0a6fd2d582d753be6628b5daecdee58acba53cba93b2bc2b5c75dee2ba48 |
cqt-plan.manifest.json |
1,721 B | 522b178e4f6e8ae5b6bf63b8e2f1a615fe2398592e27f7d9e3e219810081019f |
config.json records the I/O contract, checkpoint normalization, the CQT
configuration and the 170-label vocabulary. cqt-plan.manifest.json records the
plan's generator, configuration and payload hash.
Signature β float32 weights, opset ai.onnx 17:
| Tensor | Type | Shape | Meaning |
|---|---|---|---|
features (in) |
float32 | [16, 108, 144] |
unnormalized log(|CQT| + 1e-6) windows; 108 frames, 144 bins |
logits (out) |
float32 | [16, 108, 170] |
per-frame chord logits, before smoothing and argmax |
CQT plan β librosa 0.11.0, sr=22050, hop=2048, fmin=C1, n_bins=144,
bins_per_octave=24, norm=1, sparsity=0.01, window='hann', scale=True,
pad_mode='constant'; 6 octaves after one early downsample, 512-point FFT per
octave, resampler kaiser-lowpass-255-cutoff-0.48-beta-12.
Validation
This export + the WebGPU CQT vs the PyTorch + librosa.cqt reference runner:
| Metric | Value |
|---|---|
| per-frame chord agreement (20 instrumental stems) | 1.0000 |
| exported logits vs Torch ChordNet, identical inputs | max abs error < 1e-4 |
| degenerate outputs | 0 |
Agreement is exact because the only approximation was removed. The predecessor
artifact baked the whole pipeline into one graph with nnAudio CQT1992v2 in
place of librosa.cqt; that stand-in was the entire remaining gap (mean 0.9883,
worst 0.9410) and cost 70% of inference time and 37.8 of 47.4 MB. Reproducing
librosa's recursive per-octave transform on WebGPU fixed accuracy and size at
once.
Validate on the material fed in production β the instrumental stem. Agreement measured on audio where the reference emits a near-constant label (for example an isolated vocal, where "no chord" is correct on ~99% of frames) carries no information: a stub returning that label scores just as well. Re-run the parity gate on the exact published bytes before relying on it.
Source & lineage
Code license and weight license are separate; ONNX conversion does not change the weight license. Documented only as far as it is verifiable.
- Architecture: ChordNet "2E1D" β frequency encoder + time encoder + decoder, a small transformer (~2.3 M parameters).
- Reference implementation and weights:
ptnghia-j/ChordMini, MIT (per itsLICENSE). Upstream publishes no Hugging Face repo, so the weights come from the GitHub repository rather than the Hub. - Checkpoint:
checkpoints/2e1d_model_best.pthβ 27,523,646 B, git blobb61f6b3a02cc42b87afa38392f80d185a49f719aβ fetched at export time fromraw.githubusercontent.com. That URL tracksmainand upstream publishes no tagged release, so the fetch follows a moving branch; the blob hash above identifies what this export actually used. - Vendored code: the inference subset lives under
musetric_toolkit/chords_audio/chordminiin musetric-toolkit; see itsthirdPartyNotices.md. - Export tooling:
scripts/onnx/chordminiin musetric-toolkit. - Host runtime:
packages/cqt(the CQT) andpackages/ai(the session and the smoothing/argmax passes) inmusetric.
This export preserves the upstream MIT license; we do not claim authorship of the original weights.
License
MIT, inherited from the upstream weights.
- Downloads last month
- 160