Solitito β€” guitar chord, note and onset recognition

Models for Solitito, a real-time guitar trainer written in Rust. Recognition runs locally on the CPU, using an ordinary mono guitar signal from an audio interface or microphone.

Take7 combines chord recognition and the Rise onset detector in one self-contained ONNX file. It is the model for Solitito 0.5.7 and the current take7 source build. Earlier released binaries need updating to use its two-input contract. The model's take number is independent of the application version.

Files

File Purpose
best_model_v2_take7.onnx Current combined model: root, chord quality, sounding pitches and Rise onsets
dsp_weights.json Sparse pseudo-CQT kernel for the chord feature extractor
best_model_v2_take6_onset.onnx Previous model with the original onset head, retained for older applications and comparisons
best_model_v2_take6.onnx Previous three-output chord and pitch model

Take7 needs no separate Rise ONNX file. The DSP weights are still required.

Why a new onset detector?

A pitch detector answers which notes are sounding. Guitar practice also needs to know which notes were just struck: a string may still be ringing when the exercise moves to another chord containing the same note.

The take6 onset head used changes in CQT/chroma features and the chord encoder's context. Rise has its own short, causal spectral input and a network dedicated to detecting fresh attacks for each of the twelve pitch classes. It combines the current spectrum with its positive change relative to recent frames, so it can look for a new note while other notes continue ringing. Several pitch classes can have an onset at the same time.

In the application, Rise detections become timestamped events. With Credit only what was struck enabled, those events supply fresh-attack evidence for note practice. A consumed event cannot be reused for the next target. Chord recognition and sounding-pitch prediction retain their own branch and schedule.

Practice feedback during development reported substantially fewer repeated credits with Rise. Earlier variants also required some quiet notes to be played again. These are observations from practice, not a controlled accuracy benchmark; the historical take6 scores below do not measure the new detector.

Model inputs and outputs

The model consumes features, not raw audio. Both paths use mono audio at 16 kHz and a 256-sample feature hop.

Input Shape used by the application Features
features [1, 48, 168] 144 CQT bins, 12 chroma values and 12 bass-energy values per frame
short_features [1, 770, 35] Causal log-magnitude spectra from 1024- and 2048-sample Hann windows, through 4 kHz

The chord branch retains the CNN with Squeeze-and-Excitation blocks and a Transformer encoder with a CLS token. Its 48-frame context spans approximately 0.77 seconds.

Rise uses 64 ms and 128 ms spectral windows, a feature projection and four causal residual convolution layers with dilations 1, 2, 4 and 8. A second projection receives positive spectral growth relative to the preceding four frames. That calculation is part of the ONNX graph. The runtime supplies 34 past feature frames and the current frame; it requires no future audio.

Output Shape at these input sizes Meaning
root_logits [1, 13] Twelve pitch classes plus Noise
quality_logits [1, 11] maj, min, maj7, dom7, min7, m7b5, dim7, aug, sus, note, N
pitch_logits [1, 12] Apply sigmoid for sounding-pitch probabilities
onset_logits [1, 12, 35] Apply sigmoid; the last frame provides the current onset probabilities

Root and quality are categorical predictions; pitch and onset are independent per-class predictions. Model metadata stores the selected thresholds and onset feature contract. Use the exported onset threshold rather than assuming it is the same as the application's sounding-note threshold.

Solitito loads each independent branch into its worker's memory from the same file. Chord inference runs every 40 ms and Rise every 16 ms, without rerunning the chord encoder for each onset frame. No derived model files are written. The 16 ms update interval is not a claim of 16 ms end-to-end detection latency.

Training

The supported entry point is dist/model_trainer.py in the Solitito repository: a standalone Python script suitable for copying into a Kaggle notebook with the datasets attached.

Take7 can reuse a take6 chord checkpoint and train only Rise while keeping the chord weights frozen. The original fc_onset head is removed. Take6 checkpoints do not contain Rise weights: Rise starts from fresh weights unless an existing compatible Rise PyTorch checkpoint is supplied. Saved take7 runs can resume.

The trainer also supports training the entire model from scratch, including without Hugging Face history. Its modes are auto, onset_only, full and export_only. The last mode combines already trained branches without another training run. The exporter verifies the combined model against the original branch outputs before saving the final ONNX.

Data

  • Chord branch: synthetic guitar recordings rendered through NAM amp models, plus GuitarSet accompaniment material with corrected label handling and splits by source.
  • Rise: synthetic plucks and sustained-note mixtures, plus both solo and accompaniment recordings from GuitarSet. Synthetic cases include a held root or triad, adding a third, fifth or octave, and striking a sounding note or triad again.

GuitarSet provides note timing and pitch annotations separated by string. They are used as supervision; the detector receives mono audio, not separate string channels. An ordinary guitar does not need a hexaphonic pickup. Note-start annotations are not verified labels of picking technique.

For Rise, GuitarSet players 00–03 form the training set, player 04 is used for validation, and player 05 for testing. Synthetic groups are separated between splits. Onset targets cover 96 ms; twelve pitch classes cannot distinguish all closely spaced attacks on different strings playing the same pitch class.

Keep training_summary_v2_take7.json with the model. It records the selected threshold, source hashes, export checks and metrics for that training run. Event precision, recall, false events per minute and audio-time latency are reported separately by data domain. They are not direct measurements of the application's exercise credits.

Historical take6 results

These figures belong to take6, not Rise or a new take7 evaluation.

Chord/pitch metric Source-grouped validation, solo chord tracks excluded
Root accuracy 98.1%
Pitch F1 0.909
Exact match: root and quality 92.4%

The original take6 onset head reported F1 0.812 under its earlier evaluation protocol. This is not directly comparable with the event-level Rise metrics. Reusing a frozen chord checkpoint preserves its weights; exporting both branches in one file does not itself improve recognition accuracy.

Using take7

Use Solitito 0.5.7 or a build with take7 support. Put best_model_v2_take7.onnx and dsp_weights.json in the application's working directory. Normal startup selects take7 automatically.

./solitito --check
./solitito

For the Linux release package, ./solitito.sh sets the working directory for you; ./solitito.sh --check runs the compatibility check. On Windows, use solitito.exe --check and solitito.exe.

--check runs both branches and reports the onset threshold. SOLITITO_MODEL=/path/to/model.onnx selects a different primary file. Leave SOLITITO_ONSET_MODEL unset to use take7's own Rise branch. Neither audio recording nor SOLITITO_ONSET_RESCUE=1 is required for normal startup; the latter is an optional runtime confirmation path, disabled by default.

Custom integrations must match the feature extraction and alignment exactly, including causal padding, resampling, log scaling and feature precision. dist/gen_weights.py generates the chord DSP kernel; dist/short_onset_features.py defines the onset feature extraction. See the repository's docs/training-take7.md, docs/running.md and docs/how-it-works.md for details.

License

MIT. GuitarSet is CC BY 4.0 β€” Qingyang Xi, Rachel M. Bittner, Johan Pauwels, Xuzhou Ye and Juan P. Bello. See the GuitarSet project for attribution and dataset documentation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support