Google DeepMind released SL2T, an AI model that translates sign language directly into text, on August 12, 2026. Sign language AI, which had stayed inside the research lab, now ships inside two consumer features on Pixel 11: Gboard and Live Transcribe. The launch covers American Sign Language to English, at no additional cost.

A sign language counterpart to voice dictation

Dictating text by voice has been a standard smartphone feature for years. Turning sign language into text, by contrast, has never quite reached everyday use. There are more than 200 sign languages worldwide, used by an estimated 70 million Deaf and hard of hearing people, yet the progress that automatic translation and dictation brought to spoken languages has not carried over.

SL2T is aimed squarely at that gap. On Pixel 11, Gboard accepts sign input, so anywhere a user would previously have typed, they can now sign to the camera instead. That covers web searches, drafting messages and documents, and issuing requests to Gemini. In Live Transcribe, a user can follow the other party's speech as text and reply by signing rather than typing back and forth. Google says testers found signing in ASL faster and more natural than typing in English.

Sign language is not English on the hands

Speech recognition and sign language translation are hard in different ways. Transcribing speech maps sound to text sequentially within a single language. Sign languages are independent natural languages with their own grammars and lexicons, so the task calls for genuine machine translation rather than a word-by-word substitution.

The model also has to see movement. Sign languages carry meaning through simultaneous motion of the hands, arms, torso, head, and face, and tracking all of that at a high frame rate is a demanding computer vision problem. Google DeepMind points to early efforts such as sign language gloves as examples of what happens when a project assumes sign language is simply English expressed with the hands.

Only skeletal coordinates leave the device

SL2T never handles raw camera footage. An on-device model, MediaPipe Holistic, tracks the positions of points on the signer's body, and only that sequence of geometric coordinates is sent to the server. The original video is discarded immediately, a design choice made to protect privacy.

The translation stage carries its own design decision. Prior work on sign language translation has generally passed through an intermediate notation known as glosses. Glosses drop information specific to sign languages, including non-manual markers and spatial constructions. By translating the landmark sequence straight into text, SL2T removes artificial vocabulary limits and lets translation quality scale with the volume of data.

Training used more than 100,000 hours of data spanning over 50 sign languages, with roughly a quarter of it in ASL. Training jointly across languages, dialects, and proficiency levels produced a model that outperformed single-language versions. On the sd-test split of FLEURS-ASL, a benchmark for ASL to English translation quality, SL2T scored 70 BLEURT zero-shot, well above previously reported figures.

The unglamorous work of making it usable

A strong benchmark score and a feature people can actually use are separate things. Google DeepMind says it spent considerable effort minimizing streaming latency, preventing the model from producing text when the camera sees no signing, and keeping accuracy steady for the roughly 10 percent of signers who are left-handed. One-handed signing, which happens whenever someone is holding a phone in the other hand, received the same treatment.

Built with the Deaf community, not just for it

The project began with Sam Sepah, a Deaf Googler. Data collection ran with Deaf partners, evaluation went through Deaf user studies, and Deaf experts took part in assessing the technology's impact.

Google DeepMind also formed the AI Sign Language Advisory Committee (AISLAC), which brings together Deaf organizations and subject-matter experts worldwide, and published a joint impact report alongside the release of SL2T 1.0. The report sets out both what the technology can do and where it currently falls short, and the company plans to repeat the approach for future major sign language releases.

Support starts with the ASL and English pair, but Google DeepMind says it is working to add languages and devices, and to move into sign language generation as well.

Summary

SL2T translates sign language directly into text without passing through glosses, and it now ships inside Gboard and Live Transcribe on Pixel 11 at no extra cost. It was trained on more than 100,000 hours of data across over 50 sign languages and scored 70 BLEURT zero-shot on the sd-test split of FLEURS-ASL. Sending only on-device skeletal coordinates, and accounting for left-handed and one-handed signing, points to a model built for real use rather than for benchmarks alone. Coverage is limited to ASL into English on a single phone family, but this is the first time sign language AI has shipped as a product.