KRAFTON's flagship speech LM pair, the heart of the Raon launch: Raon-Speech-9B (TTS/STT/SpeechChat/TextQA in one end-to-end model) claims #1 among public sub-10B speech models for both English and Korean across 40 benchmarks; Raon-SpeechChat-9B is Korea's first announced full-duplex dialogue model, responding while the user is still speaking. CC-BY-NC-4.0.

Outputs 3

Raon-Speech-9B

model

End-to-end 9B speech LM "built on Qwen3 (36 layers, 4096 hidden dim)" per the model card, with a Qwen3OmniMoeAudioEncoder, Mimi codec, and ECAPA-TDNN speaker encoder — the language backbone is adapted, not pretrained in-house.

Architecture DENSE
License CC-BY-NC-4.0
Base model qwen3

Raon-SpeechChat-9B

model

Real-time full-duplex variant for overlapping-speech conversation.

Architecture DENSE
License CC-BY-NC-4.0
Base model qwen3

Raon-Speech Technical Report

paper
audioopen-weightmultimodal

Related