Voice interface · Personal research
On-device, multilingual voice interface
A voice assistant that runs entirely on my own machine: a wake word model trained on 'Hey Sam', speech recognition fine-tuned on how I actually talk, and a local language model. The design problem I care about most is language. I switch between Turkish, English and Italian, and SAM follows the switch instead of making me pick a language in settings. In September 2026 it was approved for a 12-month R&D grant and incubation at Haliç AI TEKMER; I deferred it to study full time.
- Designer & developer
- Solo
- Independent · Ongoing
- R&D grant approved · deferred for MSc

01The problem
The wake word worked from the start: a model trained on 'Hey Sam' itself, not a generic phrase borrowed off the shelf. What happened right after it woke up did not. The speech recognition component had not been fine-tuned yet, so it missed a lot of what I said, and it had no idea what to do when a sentence was not in one language. Switching from Turkish to English, or into Italian while I was learning it, meant opening a settings menu and flipping a language toggle by hand, and even then recognition misfired more often than it should have.
02Research & rationale
There was no formal study, only daily use as the test. I talk to SAM the way I actually talk, in whichever language fits the moment, sometimes switching mid-session, and every mismatch between what I said and what it understood was a data point. Research on voice interfaces in everyday use shows people constantly repairing and rephrasing when a system mishears them (Porcheron et al., 2018; Myers et al., 2018); forcing a bilingual speaker to announce their language first adds a repair step before anything has even gone wrong.
03Iterations
Fine-tuning recognition and removing the manual language switch
The first version needed a person to open settings and pick Turkish or English before speaking, and still misheard often once the right language was selected. I fine-tuned the speech recognition against my own real usage instead of a generic benchmark, and replaced the toggle with automatic language detection: speak English and SAM answers in English, switch to Turkish mid-conversation and it follows, and it now handles Italian as well.
04Result
Talking to SAM no longer means managing a settings menu before it can understand you. The language you speak is the language it replies in, decided per sentence rather than per session, and recognition is tuned against how I actually talk rather than a clean lab recording. An evaluation board at Haliç AI TEKMER and KOSGEB approved it for a 12-month R&D grant and incubation in September 2026, which I have deferred for graduate study.
- Python
- OpenWakeWord
- Local LLM
- Speech recognition
05Limitations
- n = 1. SAM is tuned and tested on one speaker, me. It may work worse for everyone else, and I have no data on that yet.
- I have not reported recognition accuracy before and after fine-tuning; improvement was judged by daily use, not measured.
06What I would do differently
Code-switching is the part of SAM that turned into a real research question: what does it cost a bilingual person, in errors and in trust, when a system cannot follow a switch mid-sentence? It is the first of my research directions, and the one I would most like to study with other speakers.
07References
- Porcheron, M., Fischer, J. E., Reeves, S., & Sharples, S. (2018). Voice interfaces in everyday life. Proceedings of CHI 2018. ACM. doi:10.1145/3173574.3174214
- Myers, C., Furqan, A., Nebolsky, J., Caro, K., & Zhu, J. (2018). Patterns for how users overcome obstacles in voice user interfaces. Proceedings of CHI 2018. ACM. doi:10.1145/3173574.3173580