alt.hn

7/26/2026 at 12:36:30 AM

Inflect-Micro-v2: complete voice in 9.36M parameters

https://huggingface.co/owensong/Inflect-Micro-v2

by nateb2022

7/26/2026 at 5:48:58 AM

This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd

thanks for shearing!

by modinfo

7/26/2026 at 3:21:08 AM

Couple highlights:

> Complete local text-to-waveform speech synthesis under 10M parameters.

In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

> English only, with one fixed male voice. This is not zero-shot voice cloning.

(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

by yjftsjthsd-h

7/26/2026 at 4:33:00 PM

When would “text-to-waveform speech synthesis” ever imply speech to text?

by semiquaver

7/26/2026 at 6:42:05 PM

The HN title is "Inflect-Micro-v2: complete voice in 9.36M parameters".

by yjftsjthsd-h

7/26/2026 at 6:43:58 PM

How heavy in inference on this? The model would easily fit on many microcontroller modules, I wonder if they could run it?

by K0balt

7/26/2026 at 7:43:27 AM

The inflections are weird but this doesn't sound like a robot. Not bad!

by NetOpWibby

7/26/2026 at 5:40:59 PM

I think we need more neurons in the human brain than parameters in this model for speech. I wonder what it says about the human brain vs LLM efficiency.

by da-x

7/26/2026 at 2:06:51 PM

I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?

by billdueber

7/26/2026 at 3:34:26 PM

I use STT/TTS to interface with a local LLM for Home Assistant in my house.

by eightysixfour

7/26/2026 at 6:41:01 PM

I'm do so as well, i tried qwen3 omni 3 but it was ridiculously stupid, and i ended up with stt thinker and tts. kokoro for now

by _davide_

7/26/2026 at 2:01:19 AM

This is impressive. I wish there were a voice clone option.

by tmaly

7/26/2026 at 4:19:01 AM

With so few parameters, I imagine a voice fine-tune might be readily tractable.

by fastball

7/26/2026 at 4:39:13 PM

this is extremely encouraging for individuals/small companies being able to train pareto-frontier TTS models (specifically compute required to run vs quality of model output)

by sudb

7/26/2026 at 12:48:44 PM

I'd love to hear it but it seems your quota is exhausted.

by StilesCrisis

7/26/2026 at 2:46:28 PM

wrote a web-demo, runs in your browser

https://inflect-tts.geronimo-labs.com

code: https://github.com/geronimi73/inflect-tts

by g58892881

7/26/2026 at 4:44:56 PM

Nice! Strangely, "Nano" sounds a lot better than "Micro" to me.

On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?

by StilesCrisis

7/26/2026 at 9:21:01 PM

also, i double checked, nano is nano and micro is micro. didnt fuck that up

by g58892881

7/26/2026 at 6:18:50 PM

right. happens on my 13 too. memory leak confirmed, not sure yet what's causing it

by g58892881

7/26/2026 at 9:20:28 PM

unable to fix it. latest dev ort, dispoing tensors, restarting the session/worker. nothing, crashes every time.

defaulting to wasm on ios devices now

by g58892881

7/26/2026 at 6:29:41 AM

Amazing quality for small size, but definitely not that enjoyable to listen to.

IMHO, its at about the same quality level of historic TTS tools.

by itake

7/26/2026 at 7:18:04 AM

I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.

by stavros

7/26/2026 at 8:17:40 AM

I compared the macos Samantha just now and I guess the inflect-micro is marginally better...

by itake

7/26/2026 at 7:53:49 AM

Ivona „Joey“, „Amy“

by leobg

7/26/2026 at 4:22:35 PM

amazing! was looking for something similar

by phoenixranger

7/26/2026 at 2:55:44 AM

amazing quality for such small size!

by jsomedon

7/26/2026 at 8:55:51 AM

Alternative title: Text to speech in 9.36M, English only.

by mcbetz

7/26/2026 at 4:05:13 PM

[flagged]

by fintuner

7/26/2026 at 5:49:59 PM

[dead]

by amelius

7/26/2026 at 4:48:10 PM

[flagged]

by zenith605

7/26/2026 at 3:37:05 PM

[dead]

by afdsaifdoi