Google states this is the first time such a model has been integrated into a real-world consumer product, aiming to serve the estimated 70 million people globally who are deaf or hard of hearing.
This development addresses a significant gap in AI infrastructure, as speech-to-text tools have historically ignored the unique grammar and spatial structures of visual languages.
By allowing users to sign directly into a camera to perform web searches, draft emails, or prompt the Gemini AI assistant, the model provides an alternative to manual typing.
To protect privacy, Google utilizes an on-device computer vision tool called MediaPipe Holistic to track body coordinates rather than uploading raw video, which the company claims makes the translation process both faster and more secure.
While the current release focuses on ASL, the model was trained on 100,000 hours of data spanning 50 different sign languages to identify shared patterns.
The system is designed to recognize one-handed signing and includes mechanisms to prevent "hallucinations," or the generation of false text from non-signing movements.
Google has established an advisory committee with deaf organizations to guide the technology’s deployment and plans to expand the model to support additional global sign languages in the future.