What is Inkling-Small?

Inkling-Small is an efficient multimodal AI model built to deliver strong reasoning and coding performance at a fraction of Inkling’s size. It is a Mixture-of-Experts transformer with 276 billion total parameters and 12 billion active parameters. The model was trained on NVIDIA GB300 NVL72 systems and is designed to combine high capability with more efficient inference. Inkling-Small supports native reasoning across text, images, and audio, allowing it to work across multimodal tasks without relying on separate encoders. Its context window supports up to one million tokens, making it useful for long-form reasoning, large-scale code understanding, document analysis, and agentic workflows. Users can adjust reasoning effort from minimal to extra high depending on whether they need faster responses or deeper computation. The model’s training process includes improved pre-training data, post-training with on-policy distillation from Inkling, and extended agentic coding reinforcement learning. These techniques helped Inkling-Small outperform its larger counterpart on reasoning and coding benchmarks. The model performs well in coding and tool-use harnesses and exceeds 80% on SWE-bench Verified. Its encoder-free architecture processes audio as dMel spectrograms and images as 40-by-40-pixel patches alongside text tokens. By combining efficient MoE design, one-million-token context, adjustable reasoning effort, multimodal processing, coding strength, and tool-use performance, Inkling-Small is designed for developers and teams that need capable AI with lower active compute requirements.

Pricing

Price Starts At:
$0.30 per million input tokens
Price Overview:
$0.30 per million input tokens and $1.20 per million output tokens

Integrations

Offers API?:
Yes, Inkling-Small provides an API

Screenshots and Video

Inkling-Small Screenshot 1

Company Facts

Company Name:
Thinking Machines Lab
Date Founded:
2025
Company Location:
United States
Company Website:
thinkingmachines.ai/news/inkling-small/

Product Details

Deployment
SaaS
Training Options
Documentation Hub
Support
Web-Based Support

Product Details

Target Company Sizes
Individual
1-10
11-50
51-200
201-500
501-1000
1001-5000
5001-10000
10001+
Target Organization Types
Mid Size Business
Small Business
Enterprise
Freelance
Nonprofit
Government
Startup
Supported Languages
English

Inkling-Small Categories and Features

Inkling-Small Customer Reviews

Write a Review
  • Reviewer Name: A Verified Reviewer
    Position: Developer
    Has used product for: Less than 6 months
    Uses the product: Daily
    Org Size (# of Employees): 26 - 99
    Feature Set
    Cost
    Would you Recommend to Others?
    1 2 3 4 5 6 7 8 9 10

    Great new small model

    Date: Jul 31 2026
    Summary

    Five stars from me. Inkling-Small feels like a very strong option for developers who want open-weight flexibility without jumping straight to the largest and most expensive frontier models.

    It may not be the absolute top model for every hard reasoning task, but that is not really the point. For developers building practical AI products, agents, and multimodal workflows, Inkling-Small looks like one of the most useful new open models to watch.

    Positive

    Inkling-Small is really interesting from a developer’s point of view because it hits a sweet spot between serious model capability and practical deployability. A 276B-parameter model with only 12B active parameters per token is exactly the kind of architecture that makes sense if you care about cost, speed, and scaling real AI workflows.

    I also like that it is open weights under Apache 2.0. That makes it way more appealing for developers who want to fine-tune, inspect, customize, or build on top of the model without being completely locked into a closed API.

    The multimodal support is a big plus too. Being able to work with text, images, and audio inputs gives Inkling-Small a lot of room for developer tools, coding agents, support bots, document workflows, and internal automation.

    Negative

    The main downside is that “small” here is still not tiny. Even with only 12B active parameters, this is still a large open model that will require real infrastructure if you want to host it yourself.

    I would also want to test it deeply before making it the backbone of a production coding agent. The model card and early coverage look promising, but real developer workflows expose problems that benchmarks do not always catch: messy repos, flaky tests, weird dependencies, tool failures, and long multi-step tasks.

    Read More...
  • Previous
  • You're on page 1
  • Next