IQuest Research has released the IQuest-Q1 model to the public. This is a sparse Mixture-of-Experts (MoE) language model designed for coding, reasoning, and multi-step tool usage within agent operations. The model has an overall parameter count of approximately 320 billion, with about 15 billion parameters activated per token.
The weights and model card were published on the Hugging Face platform under the IQuestLab organization on September 28th, and the inference code is available on GitHub under the special IQuest-Q1 license. The team positions IQuest-Q1 as a foundational model for agent command-line systems.
Technical Specifications and Training
According to the model card, IQuest-Q1 features 88 transformer layers and 256 experts, of which eight are activated per token. The model utilizes a hybrid attention pattern, combining three sliding window layers and one full attention layer, and supports a context window size of 524,288 tokens. Layers supporting speculative decoding are included for multi-token prediction.
IQuest recommends using the model with SGLang or vLLM on eight Graphics Processing Units (GPUs) and running it within environments such as Claude Code or Codex. The pre-training and intermediate training process involved shifting data towards code and STEM disciplines, along with adding longer-context agent trajectories, followed by three agent-focused stages.
IQuest synthesized tasks with their corresponding environments, including real APIs, MCP servers, and executable repositories, training a single policy across multiple environments. Reinforcement learning methods enabled the creation of four expert models for different scenarios—agent user experience, multi-environment operation, long-horizon tasks, and general agent operation. These models were subsequently merged into a single student via policy-based multi-teacher distillation.
Testing Results
When tested on published benchmarks, IQuest-Q1 demonstrated the following results: 64.6 on DeepSWE v1.1, 63.0 on NL2Repo, 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 55.7 on JobBench, and 29.6 on Agents' Last Exam. A comparative table shows that the model outperforms GLM-5.3 and DeepSeek-V4-Pro on NL2Repo, but lags behind DeepSeek-V4.1-Flash on DeepSWE and CyberGym tests. Results for other models are publicly available where possible.
IQuest also emphasized that the model was developed under human supervision. In one instance described in the IQuest blog, IQuest-Q1, operating in Claude Code, discovered that the slowing of the reward curve was caused by an extraneous space introduced during text decoding, which resulted in only the last step of multi-step trajectories remaining in the training loss function. After fixing this error, the average reward recovered. Researchers maintained control over the research direction, expensive experiments, and selection of versions for deployment. The team warns that the model is textual, is in its early stages, and requires continuous human oversight when performing real-world command-line tasks.

