Beating AI News Flash: Large model inference/deployment framework SGLang has added a native decision API, allowing off-the-shelf LLMs like Qwen3.8-27B and multimodal models to directly perform selection, judgment, and scoring without changing weights or retraining. Developers provide the current situation and several candidate answers, and the model directly returns the probability of each option, instead of first generating a block of text and then parsing it.
It leverages the "next token probability" that large models already compute. For example, if the candidate answers are A, B, and C, a regular chat model would continue generating text; SGLang instead directly reads the model's probabilities for A, B, and C to see which answer it favors. The model itself is unchanged; only the way results are read has changed.
SGLang used Qwen3.8-27B to create a "Pokémon FireRed" demo. The model decides whether to attack, switch Pokémon, or heal based on real-time game state, with a single decision taking under 100ms, and it defeated the Elite Four and Champion in one run. Qwen3.8-27B itself supports image input, so this method can also handle multimodal decision-making.
SGLang has also added a /v1/systemone interface, compatible with the TypeSafe SDK used by Jev. Applications already using this SDK can change the service address to their own SGLang instance and continue calling it. The new feature has entered the nightly version and is planned for official release with v0.5.21.

