Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
DedeProGames 
posted an update 3 days ago
Post
1480
🚀 Introducing the GRM-3.2 Family

The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.

GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.

GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.

GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.

All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many steps—whether on a server, a local workstation, or an edge device.

Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf

Organization:
OrionLLM

The comparison that would make this family legible is missing from all three cards, and it is sitting in your own base_model field.

Cliff is post-trained from Ornith-1.0-9B. Your comparison table lists GRM-2.5-Plus, GPT-5.6-Luna, Sonnet 5 and Gemini 3 Pro. It does not list the model Cliff was built on. That card publishes the same four General Agent rows:

SWE-bench Verified 69.4 against your 70.3
SWE-bench Pro 42.9 against your 43.4
NL2Repo 27.2 against your 28.5
Terminal-Bench 2.1 43.1 against your 45.3

So +0.9, +0.5, +1.3, +2.2. Sky against Ornith-1.0-35B is a much better story: +5.8 Verified, +7.9 Pro, +1.0 NL2Repo, +2.1 Terminal-Bench. Leaving the base out hides your strongest result and your weakest one at the same time.

The Terminal-Bench row is the one I would not quote yet. Ornith reports that benchmark twice, once under Terminus-2 and once under Claude Code: 43.1 and 40.6 on the 9B. That is a 2.5 point spread from scaffold choice alone, and Cliff's entire gain over base is 2.2. On the 35B the two give 64.2 and 62.8, spread 1.4, gain 2.1.

Your cards do carry a scaffold caveat, but it points outward. It warns that different labs use different agent scaffolds when reporting SWE-bench and Terminal-Bench. It does not say which one produced your own 45.3. Ornith names theirs on every row: Harbor/Terminus-2 or Claude Code 2.1.126, parser=json, temperature, top_p, context window, and how many runs were averaged.

Which scaffold produced Cliff's 45.3 and Sky's 66.3?

why not call it GRM-3? 2.7 to 3.2 is a bit strange