Silence period to detect end of speech (larger = less splitting on mid-sentence pauses)
Min Speech (ms)
Padding (ms)
Reduce padding to shorten audio sent to Whisper (trade-off with quality)
Format: Flat translation pairs — {"Japanese":"English"}. No categories or nested structures.
Ask ChatGPT etc.: "Generate a JSON with these rules: 1. Flat object format {"Japanese":"English"} 2. Keys in Japanese, values in English 3. Only proper nouns/terms likely to be spoken alone (exclude long sentences) 4. No categories or nesting 5. Keep it to 50-100 entries Theme: ◯◯" then paste here.
When dictionary matches, Llama is skipped → Zero Neuron cost.
Not started
Total audio duration sent to Whisper (for measuring cost reduction)