Whisper and wav2vec: how voice was unlocked

Duration 4:04

Comment une machine est-elle passée de quelques mots reconnus à la transcription de langues rares ? Retour sur deux modèles qui ont tout changé : wav2vec 2.0 de Meta, publié en juin 2020, et Whisper d'OpenAI, publié en décembre 2022. On explique le mécanisme réel — l'apprentissage auto-supervisé, l'architecture transformeur, l'intégration aux grands modèles de langage. Et ce que ces avancées permettent pour un assistant vocal d'entreprise.

Chapters

Every chapter below is clickable.

Transcript

Voice recognition dates back to the 1950s. Back then, it could only make out a few isolated words. Two recent models have overturned that limit. The first model is called wav2vec 2.0. Meta's research lab released it in June 2020. It changes the way machines learn. Before it, training a model required thousands of hours transcribed by hand. Each language demanded its own labelled corpus. That was the bottleneck. wav2vec 2.0 learns differently. It listens to raw audio, with no transcription. This is called self-supervised learning. The model discovers the structure of speech on its own. This approach unlocks rare languages, those with no labelled data. It also copes better with background noise. Strong accents throw it off far less. The second model is called Whisper. OpenAI released it in December 2022. Its method differs from wav2vec's. The result, though, converges. Whisper is built on a transformer architecture. It gives the model an understanding of sentence context. The model guesses a word from the ones around it. In 2023, Whisper was praised for one specific thing. It accurately transcribes low-resource languages. Where data is scarce, it holds up. Remember the difference. wav2vec learns on its own, without labels. Whisper relies on vast multilingual data and on context. Two roads to the same goal. A recent trend links these models to large language models. Raw transcription then gains in interpretation. The machine understands intent, not just words. This combination works in real time. Latency drops, the answer arrives almost instantly. That is what makes a voice conversation flow. Another development: edge computing. Processing happens on the device itself. The voice no longer leaves the terminal. This reduces privacy risks. Models are also becoming multimodal. They combine audio, text and image. This cross-referencing sharpens the understanding of a single situation. Meaning becomes clearer. Not everything is solved yet. Very noisy environments still hinder transcription. Underrepresented accents remain a weak point. The work goes on. The confidentiality of voice data remains sensitive. A recorded voice reveals a great deal about a person. At Koaee, our servers are in Europe and no conversation is ever stored. Europe is moving forward too. The French company Mistral is developing Voxtral, a model that both transcribes speech and produces it. Google and Deepgram offer their own. These advances are leaving the labs. They are reaching tools that companies use every day. The voice is becoming a real working interface. This is precisely what Koaee does. We design conversational assistants for businesses, and it is Voxtral that gives them their voice. These assistants speak sixteen languages. Our clients' customers express themselves in their own language. The assistant replies in that very language, without a detour. This film only skims the surface; the full article is waiting for you in the comments. If artificial intelligence interests you, subscribe to the channel. A like helps us. See you soon on koaee.ai.

Read the full article

Conversational assistants for business