Access Ito, a non-commercial voice model

Ito's weights are free for hobby, research, education and personal use. Tell us a little about your project: it helps us decide what to build next, and we answer every request for a commercial or collaboration license.

The Ito voice model is licensed under CC BY-NC-SA 4.0 together with Lokutor's terms of use (TERMS.md in this repository). Without a written license from Lokutor you may not use the weights, weights derived from them, or the audio they produce commercially, and you may not use Ito's output to train or improve a text-to-speech or voice model that is offered or used commercially. Ito's output is synthetic speech in the voice of a real (LibriTTS-R) speaker: say that it is synthetic when you share it, and never use it to impersonate or deceive. Small companies, startups, makers, schools and research groups can ask for a no-cost license at contact@lokutor.com.

Log in or Sign Up to review the conditions and access this model content.

Ito: natural-sounding streaming TTS for the ESP32-S3

Ito: natural speech from a $5 chip

Ito is English text-to-speech whose neural model is built to run entirely on an ESP32-S3 microcontroller (240 MHz dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. Phonemization runs on the host in this release: a computer turns the text into phoneme ids (espeak-ng) and sends them to the chip over serial, so on the chip Ito goes from phoneme ids to waveform, and the chip times below start at the ids. It streams: audio starts after a short first chunk (125 ms of audio), and the work before it does not grow with the sentence length. Its 3.3 M parameters fit in 3.8 MB of int8 weights.

Watch the 50 s video: lokutor-ai.github.io/ito/demo.mp4 (with sound) Β· Listen next to sanoTTS: lokutor-ai.github.io/ito Β· Code, firmware and tools: github.com/lokutor-ai/ito (GPLv3, commercial licenses available). From Lokutor, the makers of OΓ­do, speech recognition on the same chip.

Status (9 October 2026). The on-chip engine is verified on a laptop and in Espressif's QEMU emulator: the firmware's audio is bit-identical to the host build of the engine, and every sample below is that engine's exact output.

Measured on a board (ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), 8 October 2026, firmware fedbd26, female voice): with the main int8 set the first audio chunk (125 ms of audio) was computed and ready 184 ms after the phoneme ids were handed to the engine (text-to-phoneme conversion runs on the host and is not included) (183.9–184.3 ms over the three demo sentences), and the real-time factor was 0.66 on the boot calibration sentence (0.660–0.674 on the demo sentences). Gap-free playback is a separate number: the firmware plans to start playback about 246 ms after the phoneme ids arrive, and with that planned delay the three demo sentences had 0 underruns. No DAC or amplifier was attached and no audio was listened to, so the I2S/DAC output stage is not included, these are not acoustic measurements, and the underrun count is the firmware's own accounting against its playout clock. On this board, with nothing else running and I2S idle, synthesis ran faster than real time (RTF 0.66), with 0 underruns when playback starts after the planned 246 ms delay; this is a measurement on one board, not a guarantee. For the female voice, the other two weight sets were timed by hand once on 8 October with the tier serial command (one run each, not chosen by the boot calibration): main int4 set RTF about 0.65, first chunk about 180 ms; light set RTF about 0.61, first chunk about 168 ms. Update, 9 October 2026 (same board, same firmware fedbd26, still no DAC and nothing listened to). Male voice: the boot calibration chose the main int8 set (RTF 0.659 on the calibration sentence, in both of two boots), the bit-exact self-test passed (137,400 samples, identical to the host C engine), the first audio chunk was ready in 183.8–184.8 ms over the demo sentences, the demo-sentence RTF was 0.659–0.672, the firmware planned a gap-free start 248–249 ms after the phoneme ids arrive, and underruns were 0. The male voice's other two sets were timed by hand with tier (not boot-chosen; the firmware plans no start delay for hand-timed sets), demo sentences only: main int4 set first chunk 179.5–179.8 ms, RTF 0.645–0.653 (two passes); light set first chunk 167.9–168.0 ms, RTF 0.606–0.613 (one pass); the self-test passed with each set. Female voice, repeats: the same image booted four times in all (8 October and three times on 9 October, each after a USB-serial or watchdog reset, not a cold power cycle); the first chunk was ready in 183.5–184.5 ms over all demo runs, the boot-calibration RTF was 0.660–0.661, the planned start was 246–248 ms, the bit-exact self-test passed every time and underruns were 0. After a demo the PSRAM peak was 5372 KB (female) and 5380 KB (male) of 8192 KB, and internal SRAM 354 of 374 KB (high-water mark after the boot benchmark, 332 KB during the self-test, before any I2S DMA buffers; stack high-water marks were not printed). All of this comes from one board: the repeats and the male runs show repeatability on one chip, not a second board. Against the pre-board estimates (optimistic / central / pessimistic: RTF 0.43–0.47 / 0.63–0.66 / 0.97–0.99, first audio chunk 124–132 / 171–180 / 251–260 ms), the measured RTF of 0.66 is at the top of the central estimate and far from the pessimistic one (0.97–0.99), and the measured first chunk of 184 ms is slightly above the central estimate and below the pessimistic one. Still not measured: audio through a DAC and any listening to the output (no DAC was attached), a second board or a second chip revision, cold power-on boots (every boot was a reset), power draw, long-run thermal behaviour and supply sensitivity, stack high-water marks (not printed), the male voice's sound quality through the chip (its audio is bit-identical to the host engine: self-test PASS), and a planned start delay for the hand-timed sets (the firmware plans none). Still estimates: the optimistic, central and pessimistic columns, which are the pre-board predictions kept for comparison. In the boot benchmark the PSRAM-to-SRAM copy ran at 86.0 MB/s (above the 40–80 MB/s the estimates assumed), GDMA at 36.4 MB/s (slower than a plain copy), cached PSRAM reads at 89.2 MB/s. The firmware therefore stages weights with direct reads, the fastest of the four modes it benchmarks (mean RTF 0.66, GEMM 1.060 GMAC/s). The first public prebuilt image crash-looped on first boot on real hardware (a GDMA self-test heap overflow found on 8 October 2026, fixed in firmware commit fedbd26). The firmware ships three weight sets per voice and chooses between them at boot by measuring its own chunk times (see Weight sets below); on this board the boot calibration chose the main int8 set with a measured RTF of 0.66 (status OK, degraded flag 0; main int4 and light were not needed, so the boot did not time them; they were timed by hand afterwards). On 5 October the vocoder went from 256 to 192 channels (22 % fewer weights); in a blind test (#10) one listener heard no difference. An earlier version of this card (and our first GitHub README) said 130–210 ms and real time at 0.5–1 GOPS; that was too optimistic and is corrected here.

Voices

Two voices, each distilled from StyleTTS 2 conditioned on recordings of one LibriTTS-R reader (LibriTTS-R, CC BY 4.0). Each voice is a separate model with the same architecture, size and speed.

Voice Speaker PyTorch Chip (flash at 0x200000)
female (default) LibriTTS-R speaker 4970 ito_female.pt ito_female_esp32s3.bin
male LibriTTS-R speaker 5105 ito_male.pt ito_male_esp32s3.bin

The male voice is new (2 October 2026). Its chip file passes the same host and QEMU checks as the female voice (quantised vs float PESQ 4.55; firmware PCM bit-identical to the host engine). It has not been in a blind listening test yet; the results below are for the female voice.

Listen

The female voice, the chip engine's exact output (192-wide vocoder), for eight sentences Ito never saw in training. The male voice reads the same sentences on the demo page.

Sentence Ito (chip-exact)
01 Hey, are you still coming over for dinner tonight, or should I save you a plate?
02 Your package should arrive on Friday, October 9th, sometime before noon.
03 When I finally got to the station, the last train had already left, so I ended up sharing a taxi with two strangers who turned out to be surprisingly good company.
04 The pharmacist recommended an anti-inflammatory, but honestly, I'd rather try physiotherapy first.
05 Thanks so much for calling. I'll check the schedule and get back to you first thing tomorrow morning.
06 Could you grab some quinoa and Worcestershire sauce on your way home?
07 It's about 23 degrees outside, so you probably won't need a jacket.
08 I know it sounds strange, but I actually enjoy the quiet hours before everyone else wakes up.

The players stream from the public demo page, so they work before you accept the gate. The bit-exact WAVs are in samples/ in this repository. The demo page plays the same sentences from sanoTTS and from Ito's teacher.

Results

Blind listening test #9. One listener, who is an author of the accompanying paper, rated naturalness from 1 to 5, with system names hidden. Each system read four new conversational sentences. The Ito clips are the float PyTorch model (first release, 256-wide vocoder) with the per-sentence style predictor, not the chip build.

System Parameters Runs on Mean (4 clips)
Teacher: StyleTTS 2 (LibriTTS model) large GPU / laptop 4.75
Ito (as rated: float model with style predictor, 256-wide vocoder, 4.4 M; the shipped 192-wide one is 3.3 M, see #10) 4.4 M float model (the chip engine was rated in #10) 4.00
sanoTTS amy 1.45 M browser / desktop (not run on an MCU) 2.00
sanoTTS heart-nano 0.29 M MCU-sized (no published timing for this voice) 1.00

Blind listening test #10 (5 October 2026), the lighter vocoder: the same listener (an author), the same four sentences, the female voice, 16 clips with hidden system names; the teacher and the three chip builds only, no sanoTTS. Teacher 4.25; Ito with the 256-wide vocoder (the chip engine) 4.00; Ito with the 192-wide vocoder (the chip engine, shipped now) 4.00; the 192-wide vocoder with int4 weights (emulated) 4.00. He heard no difference between the three Ito systems (one listener, four clips per system).

Automatic metrics on the eight sentences above (measured with the 256-wide vocoder of the first release; on 60 held-out utterances the 192-wide one is within 0.02 UTMOS and +0.014 / +0.002 log-mel distance for female / male):

Teacher Ito sanoTTS amy sanoTTS heart-nano
UTMOS 4.49 4.46 3.98 2.07
WER, Whisper medium.en / base 0 / 0 % 0 / 0 % 0 / 1.0 % 1.0 / 1.0 %

Please read these with their limits:

  • One listener, an author of the paper and the builder of the systems, so not a neutral party, and four sentences per system, with integer scores. This is a direction, not a MOS study. With n = 4, differences under about half a point are noise. The names were hidden, but the sanoTTS voices differ audibly from Ito's in speaker and recording character, so "blind" does not mean unidentifiable. sanoTTS amy and heart-nano are released voices, not the 567 K voice that runs on the chip.
  • A public listening test has been prepared; it has no responses yet, and nothing from it is reported here.
  • The rated Ito clips are the float model through the same streaming path as the chip. The quantised chip engine scores PESQ 4.5 against it; test #10 below rated the chip engine itself.
  • UTMOS cannot hear intonation. Treat it as a check, not a verdict.

What we think this supports, and no more: to our knowledge, and among the systems we could find as of October 2026, the first neural TTS that streams on the ESP32-S3, and, among the neural systems built for or run on a microcontroller that we could benchmark, the highest UTMOSv2 (Ito 3.21 against 3.08 for Inflect Nano v2 by Owen Song; paired difference +0.13, 95% interval 0.03 to 0.23), a tie on UTMOS22 (4.43 vs 4.41), a lower DNSMOS (3.34 vs 3.40) and a lower word error rate (0.6 vs 1.1 %). Nine larger models, including the teacher, score higher on UTMOSv2. Inflect Nano has a third-party ESP32-P4 port (about 3.5 times slower than real time, not re-measured by us); Ito's speed is measured on one board, with the limits above. Three systems could not be benchmarked (the 567 K-parameter sanoTTS voice that runs on the chip, Moonshine Micro and Grovety tinyTTS), so this ranking covers the systems in the table only. The quality claims rest on automatic predictors and one listener who is an author. Ito is not the first neural TTS on a microcontroller and is far from the smallest.

Benchmark against other small and embedded TTS systems (54 prompts that are not in Ito's training text, identical post-processing, Whisper large-v3 word error rate; higher is better except WER; confidence intervals and paired differences in bench/ on GitHub). Ito is the chip engine's exact output with the main int8 weight set; the 256-wide row is the first release.

System Params Runs on a microcontroller UTMOSv2 UTMOS22 DNSMOS WER %
Built for or run on microcontrollers
Ito, shipped, female (chip-exact, ESP32-S3 emulated) 3.34 M (2.99 M on chip) ESP32-S3 (board; streams) 3.21 4.43 3.34 0.6
Ito, shipped, male 3.34 M (2.99 M on chip) ESP32-S3 (board; streams) 3.26 4.41 3.43 0.7
Ito, first release (256-wide vocoder) 4.05 M ESP32-S3 (emulated; streams) 3.15 4.44 3.39 0.4
Inflect Nano v2 (Owen Song) 3.97 M ESP32-P4, 3.5x slower than real time (third-party figure) 3.08 4.41 3.40 1.1
TinyTTS 1.6 M ESP32-S3, 22.9x slower than real time (per its author) 2.45 3.66 3.29 6.8
sanoTTS amy 1.45 M no (the MCU voice is a different 567 K one) 2.80 3.96 3.18 1.2
sanoTTS heart-nano 0.29 M MCU-sized, no published timing 1.33 2.17 2.97 1.7
eSpeak NG (rules) - community ports 1.74 2.14 2.76 0.3
Larger models (CPU or GPU)
Inflect Micro v2 (Owen Song) 9.36 M no 3.46 4.41 3.38 1.0
Kitten TTS nano 14.0 M no 1.99 3.93 3.31 1.1
Piper amy low 15.6 M no 3.42 4.44 3.28 0.7
Piper lessac medium 15.7 M no 3.69 4.28 3.29 0.7
MeloTTS EN 51.9 M no 3.03 3.72 3.01 3.2
Supertonic 2 65.5 M no 3.62 4.44 3.35 2.4
Kokoro-82M 81.8 M no 3.87 4.49 3.41 0.8
Supertonic 3 99.2 M no 3.84 4.45 3.32 1.7
MOSS-TTS-Nano ~100 M no 3.44 4.37 3.21 1.7
Pocket TTS 110 M no 3.31 4.33 3.31 2.5
StyleTTS 2 (Ito's teacher) 191 M no 3.43 4.47 3.34 1.5

Inflect Nano v2 has 3,966,721 parameters (3.97 M; its model card says 3.96 M). The Ito rows are the host build of the chip engine (bit-identical to the board). The authors ran this benchmark on their own system. UTMOSv2 is stochastic (about 0.03 on these means), the two Ito voices are different speakers, and the shipped-model rows were scored on a different machine from the other systems' stored values (the deterministic metrics reproduced). No teacher row exists for the male voice.

Size and compute

Parameters 3.34 M: acoustic front 1.62 M + vocoder 1.72 M (was 4.40 M with a 256-wide vocoder)
Chip weights 3.81 MB main set (ito_female_esp32s3.bin): int8 mel head and vocoder, int16 pitch path; plus the fallback sets _int4 (3.20 MB) and _light (3.05 MB)
Memory peak 5.3 of 8 MB PSRAM and 354 of 374 KB internal SRAM (high-water mark after the boot benchmark; 332 KB during the self-test; before any I2S DMA buffers, which were not active in the runs) (measured on the board, ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), 8 October 2026, firmware fedbd26, female voice; on 9 October 2026 PSRAM peak 5372 KB female / 5380 KB male of 8192 KB, internal SRAM 354 KB in both); QEMU said peak 5.3 of 8 MB PSRAM, 320 of 379 KB internal SRAM
Work before the first audio (125 ms chunk) 23.0–24.3 M instructions and 3.9 MB of weights read from PSRAM, for any sentence length (exact counts from QEMU)
Time to first audio measured: 184 ms until the first audio chunk (125 ms of audio) is computed and ready after the phoneme ids are handed to the engine (text-to-phoneme conversion runs on the host and is not included) (183.9–184.3 ms over the three demo sentences of that run; main int8 set; excludes the I2S/DAC output stage: no DAC was attached and no audio was listened to, so this is not an acoustic measurement; ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), 8 October 2026, firmware fedbd26, female voice). Pre-board estimates for this set: 124–132 / 171–180 / 251–260 ms (optimistic / central / pessimistic). Further runs on the same board and firmware, 9 October 2026, main int8 set: male voice 183.8–184.8 ms over the demo sentences (two boots); female voice 183.5–184.5 ms over all demo runs of four boots (8 October and three on 9 October, all after resets, none a cold power cycle); in every case the planned gap-free start was 246–249 ms, not 184 ms
Real-time factor measured: 0.66 on the boot calibration sentence (main int8 set; demo sentences 0.660–0.674; below 1 is faster than real time; ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), 8 October 2026, firmware fedbd26, female voice). Pre-board estimates for this set: 0.43–0.47 / 0.63–0.66 / 0.97–0.99 (optimistic / central / pessimistic). Female voice, timed by hand once, not chosen by the boot calibration: main int4 set RTF about 0.65, first chunk about 180 ms; light set RTF about 0.61, first chunk about 168 ms. Further runs, 9 October 2026: female boot-calibration RTF 0.660–0.661 over four boots; male voice boot calibration 0.659 (demo sentences 0.659–0.672); male voice timed by hand, not boot-chosen: main int4 set RTF 0.645–0.653 and first chunk 179.5–179.8 ms (two passes), light set RTF 0.606–0.613 and first chunk 167.9–168.0 ms (one pass)
Start delay for gapless speech measured: about 246 ms (246–247 ms over the three demo sentences) is the delay the firmware plans before playback starts so that speech is gap-free, planned from its own calibration; with it the three demo sentences had 0 underruns (the firmware's own accounting; no audio was output or heard; main int8 set; ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), 8 October 2026, firmware fedbd26, female voice). Further runs, 9 October 2026: 246–248 ms over four female boots; male voice 248–249 ms, with 0 underruns (the firmware plans no start delay for hand-timed sets). Pre-board estimates for this set: 125–130 / 176–182 / 558–663 ms (optimistic / central / pessimistic). delay <ms> on the serial console holds playback for a chosen time
On a laptop RTF β‰ˆ 0.01 on an Apple M4 Max CPU

The timing rows above are measured on one board (ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), firmware fedbd26; female voice on 8 and 9 October 2026, male voice on 9 October 2026); the work-per-chunk figures are exact QEMU counts. Against the pre-board estimates (optimistic / central / pessimistic: RTF 0.43–0.47 / 0.63–0.66 / 0.97–0.99, first audio chunk 124–132 / 171–180 / 251–260 ms), the measured RTF of 0.66 is at the top of the central estimate and far from the pessimistic one (0.97–0.99), and the measured first chunk of 184 ms is slightly above the central estimate and below the pessimistic one. In the boot benchmark the PSRAM-to-SRAM copy ran at 86.0 MB/s (above the 40–80 MB/s the estimates assumed), GDMA at 36.4 MB/s (slower than a plain copy), cached PSRAM reads at 89.2 MB/s. The firmware therefore stages weights with direct reads, the fastest of the four modes it benchmarks (mean RTF 0.66, GEMM 1.060 GMAC/s). The firmware benchmarks itself at boot and prints BOARD_SUMMARY. If you flash a board, please open an issue with that line and the rest of the boot log.

Stability on the board. Stability on the board (9 October 2026, one board, 240 MHz, female voice, USB powered, nothing attached, I2S idle): 600 golden-sentence iterations in about 2.5 hours (each in three modes, 1,800 runs; main int8 set 300 iterations, int4 set 100, light set 100, main int8 again with I2S idle 100) gave 1,800 PASS and 0 FAIL against the host C engine, with no panic, watchdog or reboot; a 175-token demo repeated 15 times kept a first chunk of 183.9–184.6 ms and an RTF of 0.662–0.663 with no drift. A synthetic two-core lock-step stress with Ito's own PIE kernel (about 21 minutes, every call compared with a one-core reference, including the K=704, 45 KB-tile pattern reported by the OΓ­do project) gave 0 mismatches and 0 crashes. This says no corruption or crash was observed in 600 golden-sentence iterations and a 21-minute synthetic stress on one board, and no more. OΓ­do reports faults with concurrent two-core PIE kernels on this chip; Ito did not reproduce them, but our core-1 start is a loose handshake, a cycle-aligned start was not tested, and Ito's real tiles are 5–15 KB, so anyone running dual-core PIE streaming at 240 MHz should test it, on more than one board. Not covered: a second board or chip revision, cold power-on boots, supply (USB voltage not measured), power draw, temperature, the male voice, audio output (no DAC) and runs of days. Raw logs and the stress source are in the code repository.

Weight sets and the boot check

Each voice has three chip files, flashed to three partitions: the main set (ito_<voice>_esp32s3.bin, int8), the same model with int4 weights in the ConvNeXt blocks (_int4, 3.20 MB) and a light set with one block fewer, also int4 (_light, 3.05 MB). At boot the board measures the real time of every chunk of a representative sentence with each set in turn, keeps the first whose measured real-time factor is at or below 0.85, plans the playback start delay from the measured chunk times, and, if even the fastest set measures 0.95 or more, prints a warning and sets a degraded flag instead of stuttering silently.

Pre-board estimates (optimistic / central / pessimistic, whole sentences; measured rows last, from ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz)):

main int8 main int4 light
real-time factor 0.43–0.47 / 0.63–0.66 / 0.97–0.99 0.45–0.48 / 0.62–0.64 / 0.92–0.93 0.42–0.45 / 0.59–0.61 / 0.88
time to first audio 124–132 / 171–180 / 251–260 ms 127–133 / 168–174 / 236–243 ms 118–124 / 157–164 / 222–229 ms
gapless start delay 125–130 / 176–182 / 558–663 ms 127–133 / 173–179 / 481–526 ms 118–124 / 162–169 / 406–429 ms
measured, female: real-time factor 0.66 (chosen; 0.660–0.661 over four boots) about 0.65 (timed by hand, one run, not boot-chosen) about 0.61 (timed by hand, one run, not boot-chosen)
measured, female: first audio chunk ready 184 ms (calibration sentence; 183.5–184.5 ms over all demo runs of four boots) about 180 ms (timed by hand, one run) about 168 ms (timed by hand, one run)
measured, female: planned start delay for gap-free playback 246–247 ms planned (3 demo sentences, 8 October); 246–248 ms over four boots not planned for hand-timed sets not planned for hand-timed sets
measured, male: real-time factor 0.659 (chosen at boot, both boots; demo sentences 0.659–0.672) 0.645–0.653 (timed by hand, two passes, not boot-chosen) 0.606–0.613 (timed by hand, one pass, not boot-chosen)
measured, male: first audio chunk ready 183.8–184.8 ms (demo sentences, two boots) 179.5–179.8 ms (timed by hand, two passes) 167.9–168.0 ms (timed by hand, one pass)
measured, male: planned start delay for gap-free playback 248–249 ms planned (3 demo sentences, both boots) not planned for hand-timed sets not planned for hand-timed sets

int4 weights save a fifth of the PSRAM weight traffic for about 1 % more instructions; they help only where the PSRAM is the limit. The int4 sets are bit-exact against an int8-equivalent blob, and their audio scores PESQ 4.52 (female) / 4.54 (male) against the int8 main set's; the mel distance to the float model grows by about 75 %. One listener heard no difference between the int8 and the (emulated) int4 main set in blind test #10; the light set and the male voice's int4 sets have not been heard. In the pre-board pessimistic estimate no set reached 0.85. The check works by measurement on the board that runs it; on the one board we measured, the boot calibration chose the main int8 set with a measured RTF of 0.66 (status OK, degraded flag 0; main int4 and light were not needed, so the boot did not time them; they were timed by hand afterwards). It was also tested with simulated timings in QEMU and on the host.

How it works

Text becomes phonemes on the host (espeak-ng), and token ids go to the chip over serial. On the chip, a streaming acoustic front (1.62 M; a forward-only GRU predicts durations, pitch, energy and a mel spectrogram) feeds a Vocos-style vocoder (1.72 M; ConvNeXt blocks, a harmonic pitch source and an inverse STFT) that outputs 24 kHz audio in chunks (a 125 ms first chunk on the chip, then larger ones). Ito learned its voice from StyleTTS 2 speaking as a LibriTTS-R reader. The engine is new portable C99 code with the chip's integer arithmetic; the same source builds on a laptop. The training code and recipe are not public.

Use

git clone https://github.com/lokutor-ai/ito && cd ito && pip install -e .
hf download lokutor-ai/ito --include "ito_*" --local-dir models   # both voices (accept the terms first)

ito-tts "Good morning! The coffee is ready." -o hello.wav     # PyTorch, CPU is fine; female voice
ito-tts --voice male "Good morning! The coffee is ready." -o hello_male.wav   # male voice
cd esp32/host && make && cd ../..                              # the chip engine, built for your laptop
python esp32/tools/chip_wav.py "Good morning! The coffee is ready." hello_chip.wav   # chip-exact audio
esp32/tools/flash.sh /dev/ttyUSB0                              # ESP32-S3-DevKitC-1 N16R8 + I2S DAC/amp (female voice)
VOICE=male esp32/tools/flash.sh /dev/ttyUSB0                    # ... or the male voice
python esp32/tools/say.py "Hello from a five dollar chip." --port /dev/ttyUSB0
from ito import Ito
tts = Ito.load("models/ito_female.pt")                 # female voice;  Ito.load(voice="male") or "models/ito_male.pt" for the male voice
wav = tts.synthesize("Could you grab some quinoa on your way home?")   # float32 numpy, 24 kHz
for chunk in tts.stream("A sentence of any length."):                  # 100 ms chunks (the chip starts with a 125 ms one, then larger ones)
    ...

Files

  • ito_female_esp32s3.bin: the female voice's chip weights, main set (3.81 MB), for the firmware and the host build of the engine. ito_female_esp32s3_int4.bin (3.20 MB) and ito_female_esp32s3_light.bin (3.05 MB): its fallback sets for the boot check (flash offsets 0x680000 and 0xB00000).
  • ito_female.pt: the female voice's PyTorch inference checkpoint (14 MB): front, vocoder, the fixed style the chip uses, and an optional text-to-style predictor.
  • ito_male_esp32s3.bin, ito_male_esp32s3_int4.bin, ito_male_esp32s3_light.bin, ito_male.pt: the same for the male voice. The male checkpoint has no text-to-style predictor; it uses the fixed style, as the chip does.
  • samples/01.wav … samples/08.wav: the chip engine's exact output for the sentences above (24 kHz, 16-bit).
  • LICENSE-WEIGHTS, TERMS.md, NOTICE: the license, the terms of use, and third-party attributions.

Limitations

  • English only, two voices (one per weights file), one fixed speaking style on the chip.
  • Stability was observed, not proven: no corruption or crash in 600 golden-sentence iterations and a 21-minute synthetic two-core stress, on one board only (see "Stability on the board" above).
  • Speed is measured on one board (ESP32-S3-DevKitC-1 N16R8 (rev v0.2, 8 MB octal PSRAM at 80 MHz, 16 MB flash, 240 MHz), firmware fedbd26; female voice on 8 and 9 October 2026, male voice on 9 October 2026; the repeats and the male runs are the same board); a second board or chip revision, cold power-on boots, the DAC/audio output stage and any listening, and power draw are not measured, the int4 and light sets were timed by hand only, and the optimistic / central / pessimistic figures remain pre-board estimates.
  • The pitch range is slightly narrower than the teacher's (female 0.94, male 0.97).
  • Needs an ESP32-S3 with 8 MB PSRAM (N16R8 recommended). Text normalization and text-to-phoneme conversion (espeak-ng) run on the host and the chip receives phoneme ids over serial: the 184 ms first-chunk time starts at the ids and excludes them, whereas sanoTTS runs its front end on the chip, so the two are not directly comparable. Ito is not yet a standalone text-in device.
  • The real-time factor of 0.66 was measured with both cores and about two thirds of the PSRAM in use, nothing else running and I2S idle; co-residency and I2S DMA load are not measured. The soak's self-test also times the slower single-core and 16-bit test modes (RTF up to 0.90); the shipped dual-core 8-bit mode measured 0.66. Only the main int8 set has been chosen by the boot calibration on the board.
  • Quality evidence: automatic predictors (none validated against listeners for these systems), 54 prompts, one voice and one seed per system, and one listener who is an author. The male voice's chip build has not been rated blind.
  • Disclosures. Ito is developed by Lokutor (Hashing Works, S.L.L.), which offers commercial licenses for it. The authors built the system, ran the benchmark and the listening tests, and made the board measurements; no independent party has reproduced them. Parts of the engine code, the experiment scripts and the accompanying manuscript were drafted with the help of an AI assistant (Claude, Anthropic); the authors directed the work, checked the results and the text, and take full responsibility for the content.

License and attribution

The weights and the Ito audio are licensed CC BY-NC-SA 4.0 together with Lokutor's terms of use. Commercial use needs a written license from Lokutor. That covers products and devices, paid services and APIs, internal business use and commercial content, and also weights derived from Ito. The terms also exclude training commercial TTS or voice models on Ito's output. The engine, firmware and Python package are GPLv3, with commercial licenses available.

Ito learned its voice from StyleTTS 2 (MIT) conditioned on LibriTTS-R speakers 4970 (female) and 5105 (male) (LibriTTS-R, Koizumi et al. 2023, CC BY 4.0; derived from LibriTTS and LibriVox public-domain recordings). Ito is not affiliated with or endorsed by those speakers, LibriVox or Google. No third-party weights or recordings are included. Its output is synthetic speech: please say so when you share it. Full attributions are in NOTICE.

Attribution: "Ito voice model by Lokutor (lokutor.com), CC BY-NC-SA 4.0".

We'd love to hear what you build. Lokutor offers no-cost licenses to small companies, startups, makers selling small batches, schools and research groups, and has more voices, more languages and speech recognition for the same chip. Write to contact@lokutor.com.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support