Ratings and Reviews 0 Ratings

Total
ease
features
design
support

This software has no reviews. Be the first to write a review.

Write a Review

Ratings and Reviews 1 Rating

Total
features

Alternatives to Consider

  • LTX Reviews & Ratings
    182 Ratings
    Company Website
  • Checksum.ai Reviews & Ratings
    1 Rating
    Company Website
  • LM-Kit.NET Reviews & Ratings
    29 Ratings
    Company Website
  • Runpod Reviews & Ratings
    230 Ratings
    Company Website
  • Imorgon Reviews & Ratings
    5 Ratings
    Company Website
  • Qloo Reviews & Ratings
    23 Ratings
    Company Website
  • Google AI Studio Reviews & Ratings
    30 Ratings
    Company Website
  • MEXC Reviews & Ratings
    188,765 Ratings
    Company Website
  • EBizCharge Reviews & Ratings
    207 Ratings
    Company Website
  • IUX Reviews & Ratings
    930 Ratings
    Company Website

What is Mercury 2?

Mercury 2 signifies a revolutionary leap in reasoning models, particularly tailored for instantaneous voice interactions, as it can promptly respond to incoming calls. In contrast to conventional autoregressive models that often leave callers waiting in silence while they generate responses sequentially, Mercury 2 uses a diffusion large language model architecture that can produce more than 1000 tokens per second on standard NVIDIA GPUs. This extraordinary processing speed enables it to finalize a complete reasoning cycle and start speaking in a timeframe that harmonizes with the natural flow of conversation, effectively reducing the usual wait time from several seconds to around 300 milliseconds. The functionality of Mercury models revolves around converting clear text into noise, after which a traditional Transformer is trained to reverse this process and predict the original text simultaneously across all positions. By adopting a denoising strategy that processes multiple tokens concurrently, the generation process becomes more efficient, achieving speeds comparable to customized silicon on NVIDIA H100s while enhancing responsiveness in voice applications. Consequently, Mercury 2 not only improves user interactions but also establishes a new benchmark for the field of interactive voice technology, paving the way for future advancements. With its innovative design, it promises to revolutionize the way users engage with voice systems.

What is Inkling-Small?

Inkling-Small is an efficient multimodal AI model built to deliver strong reasoning and coding performance at a fraction of Inkling’s size. It is a Mixture-of-Experts transformer with 276 billion total parameters and 12 billion active parameters. The model was trained on NVIDIA GB300 NVL72 systems and is designed to combine high capability with more efficient inference. Inkling-Small supports native reasoning across text, images, and audio, allowing it to work across multimodal tasks without relying on separate encoders. Its context window supports up to one million tokens, making it useful for long-form reasoning, large-scale code understanding, document analysis, and agentic workflows. Users can adjust reasoning effort from minimal to extra high depending on whether they need faster responses or deeper computation. The model’s training process includes improved pre-training data, post-training with on-policy distillation from Inkling, and extended agentic coding reinforcement learning. These techniques helped Inkling-Small outperform its larger counterpart on reasoning and coding benchmarks. The model performs well in coding and tool-use harnesses and exceeds 80% on SWE-bench Verified. Its encoder-free architecture processes audio as dMel spectrograms and images as 40-by-40-pixel patches alongside text tokens. By combining efficient MoE design, one-million-token context, adjustable reasoning effort, multimodal processing, coding strength, and tool-use performance, Inkling-Small is designed for developers and teams that need capable AI with lower active compute requirements.

Media

Media

Integrations Supported

Cerebras
GPT-4.1
Groq
Inception Labs
LiveKit
Model Context Protocol (MCP)
OpenAI
Pipecat
Retell AI
Tinker
Vapi AI

Integrations Supported

Cerebras
GPT-4.1
Groq
Inception Labs
LiveKit
Model Context Protocol (MCP)
OpenAI
Pipecat
Retell AI
Tinker
Vapi AI

API Availability

Has API

API Availability

Has API

Pricing Information

Pricing not provided
Free Version
Free Trial Offered?

Pricing Information

$0.30 per million input tokens
$0.30 per million input tokens and $1.20 per million output tokens
Free Version
Free Trial Offered?

Supported Platforms

SaaS
Android
iPhone
iPad
Windows
Mac
On-Prem
Chromebook
Linux

Supported Platforms

SaaS
Android
iPhone
iPad
Windows
Mac
On-Prem
Chromebook
Linux

Customer Service / Support

Standard Support
24 Hour Support
Web-Based Support

Customer Service / Support

Standard Support
24 Hour Support
Web-Based Support

Training Options

Documentation Hub
Webinars
Online Training
On-Site Training

Training Options

Documentation Hub
Webinars
Online Training
On-Site Training

Company Facts

Organization Name

Inception

Company Location

United States

Company Website

www.inceptionlabs.ai/blog/mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone

Company Facts

Organization Name

Thinking Machines Lab

Date Founded

2025

Company Location

United States

Company Website

thinkingmachines.ai/news/inkling-small/

Categories and Features

Popular Alternatives

Popular Alternatives

Mercury Coder Reviews & Ratings

Mercury Coder

Inception Labs
Mercury Edit 2 Reviews & Ratings

Mercury Edit 2

Inception
Inkling Reviews & Ratings

Inkling

Thinking Machines Lab