IQuest Research releases IQuest-Q1 model: 320B MoE for agent coding with 15B active parameters
Read more
Pandaily
pandaily.com

IQuest Research releases IQuest-Q1 model: 320B MoE for agent coding with 15B active parameters

IQuest Research has released the IQuest-Q1 model to the public. This is a sparse Mixture-of-Experts (MoE) language model designed for coding, reasoning, and multi-step tool usage within agent operations. The model has an overall parameter count of approximately 320 billion, with about 15 billion parameters activated per token.

The weights and model card were published on the Hugging Face platform under the IQuestLab organization on September 28th, and the inference code is available on GitHub under the special IQuest-Q1 license. The team positions IQuest-Q1 as a foundational model for agent command-line systems.

Technical Specifications and Training

According to the model card, IQuest-Q1 features 88 transformer layers and 256 experts, of which eight are activated per token. The model utilizes a hybrid attention pattern, combining three sliding window layers and one full attention layer, and supports a context window size of 524,288 tokens. Layers supporting speculative decoding are included for multi-token prediction.

IQuest recommends using the model with SGLang or vLLM on eight Graphics Processing Units (GPUs) and running it within environments such as Claude Code or Codex. The pre-training and intermediate training process involved shifting data towards code and STEM disciplines, along with adding longer-context agent trajectories, followed by three agent-focused stages.

IQuest synthesized tasks with their corresponding environments, including real APIs, MCP servers, and executable repositories, training a single policy across multiple environments. Reinforcement learning methods enabled the creation of four expert models for different scenarios—agent user experience, multi-environment operation, long-horizon tasks, and general agent operation. These models were subsequently merged into a single student via policy-based multi-teacher distillation.

Testing Results

When tested on published benchmarks, IQuest-Q1 demonstrated the following results: 64.6 on DeepSWE v1.1, 63.0 on NL2Repo, 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 55.7 on JobBench, and 29.6 on Agents' Last Exam. A comparative table shows that the model outperforms GLM-5.3 and DeepSeek-V4-Pro on NL2Repo, but lags behind DeepSeek-V4.1-Flash on DeepSWE and CyberGym tests. Results for other models are publicly available where possible.

IQuest also emphasized that the model was developed under human supervision. In one instance described in the IQuest blog, IQuest-Q1, operating in Claude Code, discovered that the slowing of the reward curve was caused by an extraneous space introduced during text decoding, which resulted in only the last step of multi-step trajectories remaining in the training loss function. After fixing this error, the average reward recovered. Researchers maintained control over the research direction, expensive experiments, and selection of versions for deployment. The team warns that the model is textual, is in its early stages, and requires continuous human oversight when performing real-world command-line tasks.

Similar stories

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15
Read more
pandaily.com

StepFun releases preliminary version of Step 5 model: agent with 600 billion parameters and open weights expected on October 15

StepFun introduced the preliminary version of Step 5 on September 20, 2026, as its next flagship foundational model designed for long-horizon agents, according to information from Tencent Tech, DataLearner model cards, and company materials available at stepfun.com. The official English name of the model is StepFun / Step 5.

It is important to note that although the model identifier step-5-preview is already available via the product API and the open StepFun platform, the open weights in BF16 format are scheduled only for October 15, 2026, and were not available at launch. Therefore, availability via API and open weights should be considered as separate points.

The model architecture features a Mixture of Experts (MoE) sparse design with an approximate total of 600 billion parameters. Approximately 27 billion parameters are activated per token within a 92-layer 'narrow and deep' Transformer. The choice of depth is driven by agent requirements, as longer information paths through layers contribute to improved multi-step implicit reasoning and handling of long tool outputs.

The model supports a one-million-token context window, accepting both text and image input (though video input is mentioned on third-party cards). It is oriented towards applications in AI coding, software development, financial analysis, and professional agent usage. To maintain practicality with the million-token attention, sparse GQA was applied in combination with token block merging, reducing the cost of indexer and top-k selection by approximately one eighth.

During training, special attention was paid to bit-level alignment between training and inference to ensure stable MoE routing, and load-aware scheduling, speculative decoding, and FP8 paths were utilized. StepFun claims that these methods accelerate the long-horizon Reinforcement Learning (RL) process by more than three times overall.

According to data from Artificial Analysis's composite AI index, it scores 44 points. The company places this rating among leading models focused on open access. Aggregator cards also list API prices: about $1.00 per input token and $2.70 per output token per million tokens, but comparisons of metrics and cost should be viewed as statements from third parties or providers.

Until October 15, Step 5 Preview should primarily be regarded as a functional API for long-horizon agents, having only a planned commitment for open weights, rather than a release of weights on Hugging Face or a drop-in replacement for Meituan LongCat, MiniMax Code Flash, or NaiveAI OSS MoE.

Popular