DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
| Source: Apple ML Research
Tags: Apple ML Research, sign language, NLP, accessibility, EMNLP, discourse analysis, ASL
Apple researchers introduce DiscoSign at EMNLP 2026 — the first framework for discourse-aware text-to-ASL gloss translation, addressing spatial coreference, question-answer clause structures, and concept-gloss consistency that sentence-level systems miss entirely and that directly affect comprehension for Deaf and Hard-of-Hearing users.
Details
Sign language translation systems have long processed text one sentence at a time, discarding the discourse-level context that native signers depend on. Apple researchers published DiscoSign at EMNLP 2026, the first systematic computational framework for translating text into American Sign Language glosses with full discourse awareness. The paper targets three specific phenomena that sentence-level systems fail to handle: spatial coreference resolution (referenced entities must maintain consistent spatial positions throughout a conversation), Question-Answer Clause structures (pseudocleft grammatical patterns unique to ASL), and concept-gloss consistency (stable mappings between English words and ASL signs across an entire document rather than just within a sentence). Because standard translation metrics evaluate sentences in isolation, they cannot detect discourse failures. The team introduces a purpose-built suite of evaluation metrics assessing each discourse dimension separately. Experiments on both sentence-level and discourse-level datasets show DiscoSign significantly improves spatial consistency and entity tracking while maintaining competitive single-sentence translation quality — showing that discourse awareness does not degrade baseline performance. This is the first paper to systematically address these linguistic phenomena computationally, grounding the work in ASL linguistic research from Gallaudet University collaborators.