Why tone marks exist — and why reading them isn’t enough
Vietnamese is one of the few major East Asian national languages written in Latin script. Over centuries, its writing systems shifted from elite Chinese-character literacy to a romanized orthography — chữ Quốc ngữ — that made the language’s six tones visible on the page with diacritics.
That history explains a modern learner trap: you can read the tone mark and still produce the wrong pitch shape. Locals hear movement in the voice, not ink on paper. Knowing that má and mà are different words on a flashcard is not the same as matching the contour your throat must produce.
Đúng Chưa? continues the pedagogical job that writing only half-finished: make tone hearable, comparable, and correctable — one phrase at a time — without pretending a transcript or a chatbot can pass or fail your pitch.
- Tone is the word. A wrong tone is not a cosmetic accent; it is often a different word.
- Visibility ≠ production. Diacritics help you see the target; practice must train the acoustic shape.
- Honest limits. We do not sell colonial or missionary origin stories as branding — only this product truth: reading tones is not enough to speak them.
Recording check before tone feedback
Before any tone judgment, Đúng Chưa? runs a recording check — whether the audio is clear enough to evaluate. Noisy, clipped, or unstable audio gets a retry prompt — no fake feedback that would train the wrong habit.
What is Measured
When a recording is clear enough, the practice system can inspect selected features of its pitch contour rather than relying on transcription alone. Those features are compared with the controlled reference for that phrase. The trace remains a practice aid, not perfect ground truth.
Operational Parameters
- Contour direction: Whether the usable pitch trace stays relatively level, rises, falls, or changes direction over the selected phrase.
- Timing: Where a movement appears relative to the controlled reference, with uncertainty when the signal is unstable.
- Recording conditions: Whether the attempt is long, audible, and unclipped enough to continue. A failed check produces a retry, not tone feedback.
What is NOT Measured
The practice loop does not assess grammar, fluency, health, or a learner’s anatomy. Loudness helps determine whether a recording can be checked; it is not evidence that a tone is correct.
Signal vs. AI Role
A key technical and ethical boundary is established between mathematical signal matching and language intelligence helpers. AI is never the judge of your speech attempt.
Checks before comparing
The recording check determines whether input is usable. Only then may selected contour features be compared with the controlled phrase reference. Unclear input is withheld from tone judgment.
May phrase one cue
A language model may help turn a bounded discrepancy into plain language. It does not independently listen, diagnose, or decide whether the tone passes.
This prevents chatbot-style hallucinated judgments. When audio is scoreable, one correction is surfaced. When it is not, the app abstains — no fake certainty.
The "One Correction" Rule
A wall of simultaneous corrections is difficult to act on. Đúng Chưa? keeps the next retry focused on one observable practice action.
One correction is a product boundary, not a claim that every other part of the attempt was correct. The learner can try the same phrase again with a single focus.
Dialect & Accent Boundaries
Vietnamese pronunciation varies substantially across Northern (Hà Nội), Central (Đà Nẵng), and Southern (TP. Hồ Chí Minh) speech. Tone shapes, merges (for example hỏi / ngã in the South), and segment realizations differ enough that a single “pan-Vietnam” score is not honest engineering.
Early access uses a controlled, dialect-scoped native reference deck. The active reference dialect is shown before practice. Matching that deck means “close to this reference,” not “your regional variety is wrong.”
- Scoped baselines only. Northern and Central corpora exist internally for research and harvesting; the learner-facing deck is limited to validated references we can defend.
- Multi-dialect is later work. Broader Northern / Central / Southern support is an explicit product capability — not silently assumed, not marketed until separately validated.
- No universal-accent claim. Quiet-room practice against a named baseline comes first. Café robustness and all-dialect scoring are out of scope for early access.
Known Technical Limitations
Calibrating vocal contours requires capturing high-quality voice signals. We communicate our pre-launch boundaries with absolute transparency:
- Ambient Background Noise: High levels of environmental noise (e.g. coffee shop crowd static, wind, or street traffic) introduce frequency anomalies that interfere with pitch contour tracing. Calibration works best in reasonably quiet environments.
- Microphone Quality: Standard smartphone microphones are optimized for voice calls, but cheap headset mics can sometimes clip frequencies, affecting trace accuracy.
- Controlled Practice Deck: Our waitlist version utilizes a curated baseline practice deck of high-frequency phrases. It does not support arbitrary free-form speech input.
- Dialect scope: Scoring is relative to the active reference dialect, not a claim of correctness for every regional Vietnamese variety.
- Tone is lexical: Pipelines that treat tone/diacritic mistakes as cosmetic spelling noise are the wrong tool for this product. We compare pitch contours when the recording check passes — not transcript-only “sounds close enough.”