Microsoft Post-Training Harness Lifts Qwen3.5-9B on SWE-bench
Microsoft has shared new post-training work that achieves a surprisingly large jump from a small model. Using just 6K training examples and modest compute, the approach lifts Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%.
What Happened
This work is part of an emerging theme in the field: using harnesses for model post-training. The Microsoft team shows that a small training set and limited compute can produce a large gain on a challenging coding benchmark.
Why It Matters
For anyone working with open-weight models, this suggests post-training with a harness can be an inexpensive way to get strong gains without huge data or compute. It is further evidence that the post-training setup matters as much as the base model itself.
Comments
Join the conversation — sign in to comment.
No comments yet — start the conversation!