ChordMini ChordNet (2E1D) β€” classifier + CQT plan (ONNX / WebGPU)

ONNX export of the ChordMini chord recognizer (ChordNet "2E1D", 170-class large vocabulary), packaged for the musetric packages/ai runtime (onnxruntime-web on WebGPU).

The graph is the classifier only: it takes log-CQT feature windows and returns per-frame chord logits. Feature extraction is deliberately not baked in β€” the host computes a recursive constant-Q transform on WebGPU and hands the result over as a GPU buffer, so no features cross back to the CPU. This is not a drop-in audio -> chords model.

mono PCM @ 22050 Hz              (arithmetic-mean downmix β€” see Limitations)
  -> WebGPU recursive CQT        -> log(|CQT| + 1e-6) features [T, 144]
  -> pad/window, groups of 16    -> [16, 108, 144] per run
  -> chordnet.onnx               -> logits [16, 108, 170] per run
  -> WebGPU smoothing + argmax   -> chord indices [T]

cqt-plan.bin ships with the model because it defines the features the graph expects: the octave schedule, the sparse per-octave FFT basis and the resampling FIR, baked from librosa 0.11.0. Model and plan are a matched pair β€” a release therefore carries a hashable feature-extraction contract instead of an implicit one.

Normalization ((x - mean) / (std + 1e-8)) is inside the graph. CQT, windowing, smoothing and argmax stay in the host so their GPU buffers stay reusable.

Intended uses & limitations

Intended:

  • Chord recognition over music, as a stage in an audio pipeline.
  • Client/edge inference via WebGPU through onnxruntime-web.

Out of scope:

  • Standalone use without a host that computes librosa-equivalent log-CQT features, windows them to 108 frames, and applies smoothing + argmax to the logits (see musetric packages/ai and packages/cqt).
  • Use in other training frameworks β€” this is an inference-only export.

Limitations:

  • The graph is static at 16 windows of 108 frames and 144 bins. Feed a track in groups of 16 windows, pad the last group with zero windows and trim its logits back. Windows never interact inside the model, so padding changes no real logit. Any other batch fails inside the first rewritten Transpose.
  • The features must be librosa-equivalent. Substituting a different CQT is not free: an nnAudio CQT1992v2 stand-in correlates at ~0.998 yet still costs ~1.2% of frames end to end. Use the shipped plan.
  • The model is gain-sensitive. It was trained on librosa.load's arithmetic mean downmix (L+R)/2. ffmpeg -ac 1 uses an energy-preserving rematrix (L+R)/sqrt(2), i.e. a factor of √2, which log(|CQT| + 1e-6) turns into a constant log(√2) = 0.347 offset on every feature β€” after std = 1.719 a uniform +0.20 shift, enough to flip frames near a decision boundary. Downmix as the arithmetic mean.
  • Its idx_to_chord checkpoint map differs from the reference runner's idx2voca_chord() on 70 of 170 indices, in enharmonic spelling only (Db:min vs C#:min). config.json ships the runner's vocabulary.
  • Training-data provenance of the upstream checkpoint is not documented here.

How to use

The session runs the classifier; the host supplies features and consumes logits.

import * as ort from 'onnxruntime-web/webgpu';
import { createCqt, verifyCqtPlanArtifact } from '@musetric/cqt/gpu';

const session = await ort.InferenceSession.create('chordnet.onnx', {
  executionProviders: ['webgpu'],
  preferredOutputLocation: { logits: 'gpu-buffer' },
});
const device = await ort.env.webgpu.device;

// cqt-plan.bin; verifies the payload against the SHA-256 it carries.
const plan = await verifyCqtPlanArtifact(new Uint8Array(planBytes));
const cqt = createCqt(device).get({ input: pcm, output: features, sampleCount, plan });
// cqt.run(encoder) writes log features [T, 144]; pad T up to a multiple of 108.

// Copy each group of 16 windows into `groupFeatures` and run it.
const input = ort.Tensor.fromGpuBuffer(groupFeatures, {
  dataType: 'float32',
  dims: [16, 108, 144],
});
const { logits } = await session.run({ features: input });
// logits: float32 [16, 108, 170] per group; concatenate the groups, drop the
// padding windows -> uniform 9-frame smoothing -> argmax -> indices

See the musetric packages/ai host code for the full CQT, smoothing/argmax and segment-grouping pipeline.

Known defect: onnxruntime WebGPU on Adreno 6xx, and the static rewrite (2026-09-11)

Strict on-device verification (one recorded input, wasm EP of the same build as the bit-exact reference, comparator validated by negative controls) showed the WebGPU EP returns silently wrong logits for the original graph on Adreno 660 (OnePlus 9RT): 95.7% of chord logits mismatched the reference (maxAbs 0.58, nondeterministic across runs of one session). The mechanism is the shared-tile Transpose kernel of ORT's WebGPU EP: this graph contains 111 Transpose nodes, most of them (1,0,2) over tensors whose leading window axis collapses to rank-2-swapped - exactly the broken kernel family (one-line ORT fix pending upstream, transpose.cc tile + 1 -> + 8). Desktop NVIDIA is bit-exact on the same input and build.

chordnet.onnx now ships the static rewrite: all 79 affected Transposes become Reshape -> Gather(const int32 idx) -> Reshape (exact permutations of values), everything else is untouched. Verified on 9RT WebGPU against the wasm EP of the same file: 0 mismatched logits (maxAbs 1e-5), deterministic, in both storageBufferCacheMode: 'simple' and the default; faithful to the original graph within 4e-6 (verified in ORT CPU). Full evidence: the musetric plan gpu/adreno660-strict-2026-09-11.md.

That rewrite is static: every Gather index and Reshape target is baked for the captured batch. The first revision was built for one window but still declared a dynamic W, so any host feeding a whole track failed and the only working feed was one window per run, where the fixed cost of each WebGPU run dominates. This revision is built for 16 windows and declares exactly [16, 108, 144]. 16 is below the cap that matters: the frequency encoder runs Softmax over [W * 108, 8, 2, 2], and onnxruntime's WebGPU Softmax writes past its output once a dispatch has more than 65535 rows, which a batch of 38 windows or more would reach.

Checked against the wasm EP of the same file, one input of 26 and 60 windows:

runtime max |Ξ”logit| chord argmax 26 windows, 1 β†’ 16 per run 60 windows
OnePlus 9RT, Adreno 660, both cache modes 5.7e-6 identical 0.58 s β†’ 0.53 s 1.32 s β†’ 1.06 s
desktop NVIDIA, D3D12 6.7e-6 identical 0.31 s β†’ 0.03 s 0.61 s β†’ 0.06 s
Galaxy S24 Ultra, Adreno 750 1.99 1 of 2808 frames differs 0.30 s β†’ 0.18 s 1.2 s β†’ 0.35 s

In ORT CPU this revision matches the original graph within 7e-6 at 1, 5, 16, 20 and 60 windows fed in padded groups, with identical argmax.

Known defect: onnxruntime WebGPU MatMul on Adreno 750 (2026-09-20)

The Galaxy S24 Ultra row above was the second defect. There the WebGPU EP computes a float32 MatMul wrongly for two shapes this graph uses, by units and deterministically, while wasm on the same device is right: a four-dimensional product with the extents of the chord attention, and a product whose output width is not a multiple of four, which the classifier [16, 108, 144] x [144, 170] has. Measured on that device against wasm, a single node of the classifier shape differs by 4.2 at 170 columns and by 1.9e-6 at 168 or 172.

chordnet.onnx now folds the two batch axes of every four-dimensional MatMul into one and pads the classifier weight to 172 columns, slicing the result back. Both rewrites are exact, and they are what makes the graph agree with wasm on that GPU. Checked against the wasm EP of the same file, 16 windows of deterministic features:

runtime previous revision this revision
Galaxy S24 Ultra, Adreno 750, Chrome 153 max |Ξ”| 3.13, 386 of 1728 frames choose another chord max |Ξ”| 5.7e-6, no frame changed
OnePlus 9RT, Adreno 660, Chrome 152 5.3e-6, 753 ms 5.7e-6, 635 ms
desktop NVIDIA, Chrome 153 5.7e-6, 933 ms 6.7e-6, 774 ms

In ORT CPU this revision matches the exported graph within 5.7e-6, the same as the previous one, with identical argmax.

The time attention on the vec4 kernel (2026-09-30)

The cause of the Adreno 750 defect is now known (musetric#891): the driver miscompiles the scalar MatMul kernel that onnxruntime's WebGPU EP generates for two- and four-dimensional products, while its three-dimensional variant comes out right, which is why the fold above works. The EP takes the scalar kernel whenever K or N is not a multiple of four, as in the time attention of this graph, [16, 8, 108, 18] x [16, 8, 18, 108] and [16, 8, 108, 108] x [16, 8, 108, 18].

This revision takes the time attention off that kernel: K and N are padded with zeros to 20, which is the vec4 kernel, and the extra output columns are sliced off. The two-row frequency attention keeps the fold, because onnxruntime runs products of eight rows or fewer on a vec4 kernel that is wrong on Adreno 660. Both are exact.

Checked against ORT CPU on the first 16 windows of the CQT features of the instrumental stem of Rxbyn, "Bad Side" (https://www.jamendo.com/track/1556580/bad-side, CC BY 3.0, from JamendoLyrics), as the musetric parity run computes them; onnxruntime-web 1.30.0 with musetric's WebGPU provider and session options, median of 5 runs with the logits read back:

runtime previous revision this revision
Galaxy S24 Ultra, Adreno 750, Chrome 153 max |Ξ”| 4.8e-6, no frame changed, 101.2 ms 4.8e-6, no frame changed, 97.8 ms
OnePlus 9RT, Adreno 660, Chrome 153 4.8e-6, no frame changed, 286.5 ms 4.8e-6, no frame changed, 283.4 ms
desktop NVIDIA RTX 3060 Laptop, Chrome 153 4.8e-6, no frame changed, 24.0 ms 4.8e-6, no frame changed, 20.7 ms

Without any MatMul rewrite the same graph differs there by 2.07 on the Adreno 750, with 94 of 1728 frames choosing another chord.

Rebuild recipe: capture_shapes.py --ops Transpose,MatMul --windows 16 --shape 16,108,144 --input features + rewrite_static_adreno.py --vec4-attention, then dispatch_rows_audit.py, in musetric-toolkit scripts/onnx/.

Files

File Size SHA256
chordnet.onnx 17,080,550 B 1faee8e1cf168300afe0f47517e76836ac200c3841b0fb430fad7dfea5ad6394
config.json 3,009 B 1f26c11ebea51ec08f12e813eb213a729fa0ecc407ac7632dfdc7bad67e65aa4
cqt-plan.bin 23,896 B c31f0a6fd2d582d753be6628b5daecdee58acba53cba93b2bc2b5c75dee2ba48
cqt-plan.manifest.json 1,721 B 522b178e4f6e8ae5b6bf63b8e2f1a615fe2398592e27f7d9e3e219810081019f

config.json records the I/O contract, checkpoint normalization, the CQT configuration and the 170-label vocabulary. cqt-plan.manifest.json records the plan's generator, configuration and payload hash.

Signature β€” float32 weights, opset ai.onnx 17:

Tensor Type Shape Meaning
features (in) float32 [16, 108, 144] unnormalized log(|CQT| + 1e-6) windows; 108 frames, 144 bins
logits (out) float32 [16, 108, 170] per-frame chord logits, before smoothing and argmax

CQT plan β€” librosa 0.11.0, sr=22050, hop=2048, fmin=C1, n_bins=144, bins_per_octave=24, norm=1, sparsity=0.01, window='hann', scale=True, pad_mode='constant'; 6 octaves after one early downsample, 512-point FFT per octave, resampler kaiser-lowpass-255-cutoff-0.48-beta-12.

Validation

This export + the WebGPU CQT vs the PyTorch + librosa.cqt reference runner:

Metric Value
per-frame chord agreement (20 instrumental stems) 1.0000
exported logits vs Torch ChordNet, identical inputs max abs error < 1e-4
degenerate outputs 0

Agreement is exact because the only approximation was removed. The predecessor artifact baked the whole pipeline into one graph with nnAudio CQT1992v2 in place of librosa.cqt; that stand-in was the entire remaining gap (mean 0.9883, worst 0.9410) and cost 70% of inference time and 37.8 of 47.4 MB. Reproducing librosa's recursive per-octave transform on WebGPU fixed accuracy and size at once.

Validate on the material fed in production β€” the instrumental stem. Agreement measured on audio where the reference emits a near-constant label (for example an isolated vocal, where "no chord" is correct on ~99% of frames) carries no information: a stub returning that label scores just as well. Re-run the parity gate on the exact published bytes before relying on it.

Source & lineage

Code license and weight license are separate; ONNX conversion does not change the weight license. Documented only as far as it is verifiable.

  • Architecture: ChordNet "2E1D" β€” frequency encoder + time encoder + decoder, a small transformer (~2.3 M parameters).
  • Reference implementation and weights: ptnghia-j/ChordMini, MIT (per its LICENSE). Upstream publishes no Hugging Face repo, so the weights come from the GitHub repository rather than the Hub.
  • Checkpoint: checkpoints/2e1d_model_best.pth β€” 27,523,646 B, git blob b61f6b3a02cc42b87afa38392f80d185a49f719a β€” fetched at export time from raw.githubusercontent.com. That URL tracks main and upstream publishes no tagged release, so the fetch follows a moving branch; the blob hash above identifies what this export actually used.
  • Vendored code: the inference subset lives under musetric_toolkit/chords_audio/chordmini in musetric-toolkit; see its thirdPartyNotices.md.
  • Export tooling: scripts/onnx/chordmini in musetric-toolkit.
  • Host runtime: packages/cqt (the CQT) and packages/ai (the session and the smoothing/argmax passes) in musetric.

This export preserves the upstream MIT license; we do not claim authorship of the original weights.

License

MIT, inherited from the upstream weights.

Downloads last month
160
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support