Hands to words.
Spell to the camera and realsign reads it back, one letter at a time, without sending anything anywhere.
24 letters of the manual alphabet, one hand. Video never leaves your device.
interpreter
Fingerspelling, one hand, one signer. The camera and the model both run in this browser.
Show your hand to the camera, then hold one letter still at a time.
Video and landmarks stay on your device. Nothing is uploaded.
letters
Start the camera, then spell a word one letter at a time.
reading
Your words appear here.
Exactly the letters that were read, in order. Nothing corrects them into a word that looks more likely, so a misread letter stays visible instead of being tidied away.
transcript
- Finished words collect here, letter by letter beside the reading.
Trained on 1,440 samples from one sitting, one hand, one camera. The 99.7% it scores on held-out frames from that sitting is an upper bound, not a promise.
What it can read
These 24 letters of the manual alphabet, held one at a time. Not J or Z, and no word signs.
- A
- B
- C
- D
- E
- F
- G
- H
- I
- J
- K
- L
- M
- N
- O
- P
- Q
- R
- S
- T
- U
- V
- W
- X
- Y
- Z
J and Z are struck through because they are movements, not shapes. J is drawn in the I handshape and Z with a pointing index, so a single frame of either is already another letter — and the normalization removes the tilt that would otherwise separate them. They need a model that reads time, which this is not.
What this does not do
Worth reading before you rely on it for anything.
One person taught it, in one sitting
The model learned from 1,440 samples recorded through this app: 60 frames of each letter, one hand, one room, one camera. It reads 99.7% of the frames held out of that sitting, and that number says almost nothing about how it will read yours.
A word list, not a language
It reads 24 handshapes of the manual alphabet, one at a time. Fingerspelling is how ASL handles names and words it has no sign for — it is not how ASL is spoken, and it is a fraction of the language.
No J, no Z, no word signs
J and Z are traced in the air rather than held, so a single frame of either is already a different letter. Word signs are movements too. Both need a model that reads a sequence, and that does not exist here yet.
Letters that look alike get confused
M, N, S and T are the same fist with the thumb in four places, and R, U and V differ by how two fingers cross. When the model is unsure it says so with a number instead of guessing confidently.
Nothing corrects the spelling
The reading is exactly the letters that landed, in order. There is no language model tidying HELLQ into HELLO, which means a misread stays visible rather than being smoothed into something that was never signed.
Facial grammar and signing space are missing
Eyebrows, mouth and head movement carry real meaning in ASL, and real signing places signs in the space around the signer. None of that is captured by 21 landmarks on one hand.