What is Inkling-Small?
Inkling-Small is an efficient multimodal AI model built to deliver strong reasoning and coding performance at a fraction of Inkling’s size. It is a Mixture-of-Experts transformer with 276 billion total parameters and 12 billion active parameters. The model was trained on NVIDIA GB300 NVL72 systems and is designed to combine high capability with more efficient inference. Inkling-Small supports native reasoning across text, images, and audio, allowing it to work across multimodal tasks without relying on separate encoders. Its context window supports up to one million tokens, making it useful for long-form reasoning, large-scale code understanding, document analysis, and agentic workflows. Users can adjust reasoning effort from minimal to extra high depending on whether they need faster responses or deeper computation. The model’s training process includes improved pre-training data, post-training with on-policy distillation from Inkling, and extended agentic coding reinforcement learning. These techniques helped Inkling-Small outperform its larger counterpart on reasoning and coding benchmarks. The model performs well in coding and tool-use harnesses and exceeds 80% on SWE-bench Verified. Its encoder-free architecture processes audio as dMel spectrograms and images as 40-by-40-pixel patches alongside text tokens. By combining efficient MoE design, one-million-token context, adjustable reasoning effort, multimodal processing, coding strength, and tool-use performance, Inkling-Small is designed for developers and teams that need capable AI with lower active compute requirements.
Pricing
Company Facts
Product Details
Product Details
Inkling-Small Categories and Features
Inkling-Small Customer Reviews
Write a Review-
Would you Recommend to Others?1 2 3 4 5 6 7 8 9 10
Great new small model
Date: Jul 31 2026SummaryFive stars from me. Inkling-Small feels like a very strong option for developers who want open-weight flexibility without jumping straight to the largest and most expensive frontier models.
It may not be the absolute top model for every hard reasoning task, but that is not really the point. For developers building practical AI products, agents, and multimodal workflows, Inkling-Small looks like one of the most useful new open models to watch.PositiveInkling-Small is really interesting from a developer’s point of view because it hits a sweet spot between serious model capability and practical deployability. A 276B-parameter model with only 12B active parameters per token is exactly the kind of architecture that makes sense if you care about cost, speed, and scaling real AI workflows.
I also like that it is open weights under Apache 2.0. That makes it way more appealing for developers who want to fine-tune, inspect, customize, or build on top of the model without being completely locked into a closed API.
The multimodal support is a big plus too. Being able to work with text, images, and audio inputs gives Inkling-Small a lot of room for developer tools, coding agents, support bots, document workflows, and internal automation.NegativeThe main downside is that “small” here is still not tiny. Even with only 12B active parameters, this is still a large open model that will require real infrastructure if you want to host it yourself.
Read More...
I would also want to test it deeply before making it the backbone of a production coding agent. The model card and early coverage look promising, but real developer workflows expose problems that benchmarks do not always catch: messy repos, flaky tests, weird dependencies, tool failures, and long multi-step tasks.
- Previous
- You're on page 1
- Next