-
1
Atlas by World Labs
World Labs
Transforming imagination into reality with unparalleled spatial intelligence.
Atlas is a cutting-edge omni world model crafted for spatial intelligence, effortlessly interacting with text, images, video, and 3D data. As a sophisticated multimodal autoregressive diffusion transformer, it harmonizes diverse inputs into a unified spatial framework, while simultaneously forecasting subsequent elements and maintaining coherence in three-dimensional spaces through its insights and creative ideas. This robust model enables world creation, reconstruction, and simulation across a wide array of applications. It is adept at generating images and videos from one or several reference images, providing meticulous camera control to produce extensive and coherent videos that adhere to custom-designed camera paths. In the realm of spatial reconstruction, Atlas shines in reinterpreting real-world settings from a limited number of input images, which allows it to create distinct viewpoints and deliver detailed 3D outputs such as point clouds and 3D Gaussian splats. With the incorporation of additional input perspectives, the model can further improve contextual understanding, resulting in progressively precise reconstructions with a diminished dependence on imagination. Additionally, Atlas has the capability to analyze and forecast the progression of worlds over time, thereby adding a dynamic aspect to its functionality, which enhances its versatility in various scenarios. This multifaceted approach not only enriches the user experience but also expands the potential applications of the model across different fields.
-
2
K2 Horizon
Institute of Foundation Models
Unleash powerful, dynamic performance across every computational task!
K2 Horizon consists of a collection of six open models, namely the 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B, each meticulously designed to excel in specific areas such as reasoning, mathematics, coding, agentic tasks, and overall functional capabilities. These models are built upon a cohesive architecture that incorporates a shared vocabulary, training methodologies, interfaces, evaluation frameworks, and deployment tools, which enable smooth transitions between different model sizes and effective management of varying workloads. Leading the lineup, the 375B-A23B model excels in complex reasoning, software development, research endeavors, and long-term agentic operations, whereas the 32B and 36B-A4B models emphasize strong local deployment capabilities. The 36B-A4B model is distinguished by its cutting-edge Mixture-of-Value Attention mechanism, which fuses sparse attention with Mixture-of-Experts layers, enabling it to utilize around 4 billion parameters per token, thereby closely rivaling the performance of the denser 32B model. This innovative architecture not only enhances versatility but also optimizes resource utilization across a diverse array of applications, making K2 Horizon a formidable presence in the field of model technology. Additionally, the collective strengths of these models allow for a comprehensive approach to tackling various challenges within their respective domains.
-
3
MAI-Transcribe-2
Microsoft AI
Revolutionize transcription with unparalleled accuracy and versatility.
MAI-Transcribe-2 stands as the apex of Microsoft AI's transcription technology, meticulously designed to deliver swift and accurate speech recognition in a variety of real-world audio settings. The model boasts functionalities such as speaker diarization, which allows it to distinguish among speakers and attribute dialogue accurately, along with providing word-level timestamps that enhance alignment, searching, navigation, and editing capabilities. It also incorporates keyword biasing to boost the recognition accuracy of specialized terminology, abbreviations, and names that might otherwise be difficult to discern in context. Developers can choose from various transcription styles, including a verbatim option that captures filler words and false starts for comprehensive analysis and compliance, or a clean option that omits these elements for clearer captions and more polished published transcripts. Moreover, the model is skilled at managing code-switching, effortlessly shifting between languages in conversations, and it can even handle mixed language combinations like Hinglish and Spanglish, all while automatically detecting the spoken language. This adaptability not only enhances usability but also positions MAI-Transcribe-2 as an indispensable asset in multilingual environments and diverse applications. Consequently, its innovative features cater to the evolving demands of transcription across industries.
-
4
Mercury 2.5
Inception
Unleash unparalleled performance with the ultimate diffusion language model.
Mercury 2.5 exemplifies the ultimate achievement in production models from Inception, showcasing a significant quality improvement over its predecessor, Mercury 2, while maintaining an outstanding low-latency performance. This model is recognized as the most sophisticated diffusion language model currently on the market and is claimed by Inception to be the largest diffusion LLM ever created. With a remarkable 40% increase in intelligence compared to Mercury 2, its capabilities closely mirror those of economical frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. It features an impressive generation rate of 1,107 tokens per second on widely used NVIDIA GPUs and can handle a substantial 260K-token context window. Among its many attributes are customizable reasoning abilities, concurrent tool executions, and JSON compatibility with schemas. Designed specifically for tasks sensitive to latency, it excels in environments that require multiple model calls during one interaction. In various applications, including search agents and RAG pipelines, Mercury 2.5 performs exceptionally well in activities such as planning, query rewriting, re-ranking, fact structuring, source summarization, and answer verification. This efficiency ensures swift response times, making it a vital asset for developers aiming to enhance their workflow productivity. As technology continues to evolve, models like Mercury 2.5 will likely set new benchmarks in the industry.
-
5
Grok 4.8
SpaceXAI
Unlock next-level coding and reasoning for AI workflows.
Grok 4.8 is an upcoming frontier AI model from xAI expected to improve reasoning, coding, agentic execution, and professional knowledge work across the Grok ecosystem. Elon Musk has described Grok 4.8 as a roughly 2.5-trillion-parameter model, representing an increase in scale over the planned 2.1-trillion-parameter Grok 4.7. The model is being trained using a new C++ training software stack rather than the infrastructure used for some earlier Grok training runs. xAI expects the initial training phase to complete before the model moves into reinforcement learning, evaluation, and additional post-training refinement. Grok 4.8 is anticipated to extend the current Grok generation’s emphasis on software development, complex reasoning, agentic tool calling, and professional knowledge tasks. Grok 4.7, the current officially documented flagship, supports image and text input, configurable reasoning effort, and a 500,000-token context window. A larger successor could provide additional capacity for difficult coding assignments, research, application development, data analysis, and long-running tasks that require sustained planning and verification. The model may also strengthen xAI products such as Grok Build and persistent AI agents that work across applications and execute multi-step jobs. Musk has characterized the developing model as an improvement over the preceding Grok generation, but independent benchmarks are not yet available to confirm its eventual performance. xAI has not announced final pricing, context length, inference speed, API naming, benchmark scores, or a general availability date for Grok 4.8. Grok 4.8 is expected to target software developers, AI engineers, researchers, enterprises, and organizations building advanced autonomous agents and computational workflows.
-
6
StepAudio 3
StepFun
Revolutionizing audio interaction: voice, sound, and music mastery.
StepAudio 3 marks a significant leap in StepFun's series of audio models, crafted to understand, generate, and interact through different sound modalities such as voice, music, and ambient noises. This series includes a range of specialized models, namely StepAudio 3 Realtime for interactive full-duplex conversations, StepAudio 3 ASR focused on precise speech recognition, StepAudio 3 TTS designed for smooth speech synthesis, StepAudio 3 Gen for diverse audio creation, and StepAudio 3 Music tailored for composing lengthy musical works. The Realtime model innovatively operates with an ongoing loop of listening, conversing, reflecting, and responding, skillfully capturing not only spoken language but also subtle cues like laughter, hesitation, emotions, and interruptions. Unlike conventional systems, it can simultaneously process information and provide replies, adeptly handle intricate inquiries without breaking the flow of dialogue, and utilize various tools to accomplish tasks once it comprehends the user’s intent. In addition, StepAudio 3 Gen merges multiple capabilities such as zero-shot TTS, voice design, vocal creation, sound effects, and mixed audio generation into a unified platform. Meanwhile, StepAudio 3 Music excels in crafting songs controlled by text, instrumental pieces, and vocal compositions, thus serving as a robust asset for audio innovation. This groundbreaking suite highlights the fusion of interaction and creativity, setting new standards for the potential of audio models in various applications. Each model contributes to a holistic experience, ensuring that users are empowered to explore a vast landscape of auditory possibilities.
-
7
BLOOM
BigScience
Unleash creativity with unparalleled multilingual text generation capabilities.
BLOOM is an autoregressive language model created to generate text in response to prompts, leveraging vast datasets and robust computational resources. As a result, it produces fluent and coherent text in 46 languages along with 13 programming languages, making its output often indistinguishable from that of human authors. In addition, BLOOM can address various text-based tasks that it hasn't explicitly been trained for, as long as they are presented as text generation prompts. This adaptability not only showcases BLOOM's versatility but also enhances its effectiveness in a multitude of writing contexts. Its capacity to engage with diverse challenges underscores its potential impact on content creation across different domains.
-
8
NVIDIA NeMo Megatron is a robust framework specifically crafted for the training and deployment of large language models (LLMs) that can encompass billions to trillions of parameters. Functioning as a key element of the NVIDIA AI platform, it offers an efficient, cost-effective, and containerized solution for building and deploying LLMs. Designed with enterprise application development in mind, this framework utilizes advanced technologies derived from NVIDIA's research, presenting a comprehensive workflow that automates the distributed processing of data, supports the training of extensive custom models such as GPT-3, T5, and multilingual T5 (mT5), and facilitates model deployment for large-scale inference tasks. The process of implementing LLMs is made effortless through the provision of validated recipes and predefined configurations that optimize both training and inference phases. Furthermore, the hyperparameter optimization tool greatly aids model customization by autonomously identifying the best hyperparameter settings, which boosts performance during training and inference across diverse distributed GPU cluster environments. This innovative approach not only conserves valuable time but also guarantees that users can attain exceptional outcomes with reduced effort and increased efficiency. Ultimately, NVIDIA NeMo Megatron represents a significant advancement in the field of artificial intelligence, empowering developers to harness the full potential of LLMs with unparalleled ease.
-
9
ALBERT
Google
Transforming language understanding through self-supervised learning innovation.
ALBERT is a groundbreaking Transformer model that employs self-supervised learning and has been pretrained on a vast array of English text. Its automated mechanisms remove the necessity for manual data labeling, allowing the model to generate both inputs and labels straight from raw text. The training of ALBERT revolves around two main objectives. The first is Masked Language Modeling (MLM), which randomly masks 15% of the words in a sentence, prompting the model to predict the missing words. This approach stands in contrast to RNNs and autoregressive models like GPT, as it allows for the capture of bidirectional representations in sentences. The second objective, Sentence Ordering Prediction (SOP), aims to ascertain the proper order of two adjacent segments of text during the pretraining process. By implementing these strategies, ALBERT significantly improves its comprehension of linguistic context and structure. This innovative architecture positions ALBERT as a strong contender in the realm of natural language processing, pushing the boundaries of what language models can achieve.
-
10
ERNIE 3.0 Titan
Baidu
Unleashing the future of language understanding and generation.
Pre-trained language models have advanced significantly, demonstrating exceptional performance in various Natural Language Processing (NLP) tasks. The remarkable features of GPT-3 illustrate that scaling these models can lead to the discovery of their immense capabilities. Recently, the introduction of a comprehensive framework called ERNIE 3.0 has allowed for the pre-training of large-scale models infused with knowledge, resulting in a model with an impressive 10 billion parameters. This version of ERNIE 3.0 has outperformed many leading models across numerous NLP challenges. In our pursuit of exploring the impact of scaling, we have created an even larger model named ERNIE 3.0 Titan, which boasts up to 260 billion parameters and is developed on the PaddlePaddle framework. Moreover, we have incorporated a self-supervised adversarial loss coupled with a controllable language modeling loss, which empowers ERNIE 3.0 Titan to generate text that is both accurate and adaptable, thus extending the limits of what these models can achieve. This innovative methodology not only improves the model's overall performance but also paves the way for new research opportunities in the fields of text generation and fine-tuning control. As the landscape of NLP continues to evolve, the advancements in these models promise to drive further breakthroughs in understanding and generating human language.
-
11
EXAONE
LG
"Transforming AI potential through expert collaboration and innovation."
EXAONE is a cutting-edge language model developed by LG AI Research, aimed at fostering "Expert AI" in multiple disciplines. To bolster EXAONE's capabilities, the Expert AI Alliance was formed, uniting leading companies from various industries for collaborative efforts. These partner organizations will serve as mentors, providing their knowledge, skills, and data to help EXAONE excel in targeted areas. Similar to a college student who has completed their general studies, EXAONE needs specialized training to achieve true mastery in specific fields. LG AI Research has already demonstrated the potential of EXAONE through real-world applications, such as Tilda, an AI human artist that premiered at New York Fashion Week, and AI tools that efficiently summarize customer service interactions and extract valuable insights from complex academic texts. This initiative underscores not only the innovative uses of AI technology but also the critical role of collaboration in pushing technological boundaries. Moreover, the ongoing partnerships within the Expert AI Alliance promise to yield even more groundbreaking advancements in the future.
-
12
Jurassic-1
AI21 Labs
Unlock creativity with the most advanced language model.
Jurassic-1 features two distinct model sizes, with the Jumbo variant being the most expansive at 178 billion parameters, showcasing the highest level of intricacy among language models available to developers. Presently, AI21 Studio is undergoing an open beta phase, encouraging users to sign up and start engaging with Jurassic-1 via a user-friendly API and an interactive online platform.
At AI21 Labs, we aim to transform the way individuals interact with reading and writing by incorporating machines as cognitive partners, a vision that necessitates collaborative efforts to achieve. Our journey into the realm of language models began during what we call our Mesozoic Era (2017 😉). Building on this initial research, Jurassic-1 represents the first series of models we are now making available for widespread public use. Looking ahead, we are eager to witness the innovative ways in which users will harness these technological advancements in their creative endeavors. Furthermore, we believe that this collaboration between humans and machines will unlock new frontiers in communication and expression.
-
13
Alpaca
Stanford Center for Research on Foundation Models (CRFM)
Unlocking accessible innovation for the future of AI dialogue.
Models designed to follow instructions, such as GPT-3.5 (text-DaVinci-003), ChatGPT, Claude, and Bing Chat, have experienced remarkable improvements in their functionalities, resulting in a notable increase in their utilization by users in various personal and professional environments. While their rising popularity and integration into everyday activities is evident, these models still face significant challenges, including the potential to spread misleading information, perpetuate detrimental stereotypes, and utilize offensive language. Addressing these pressing concerns necessitates active engagement from researchers and academics to further investigate these models. However, the pursuit of research on instruction-following models in academic circles has been complicated by the lack of accessible alternatives to proprietary systems like OpenAI’s text-DaVinci-003. To bridge this divide, we are excited to share our findings on Alpaca, an instruction-following language model that has been fine-tuned from Meta’s LLaMA 7B model, as we aim to enhance the dialogue and advancements in this domain. By shedding light on Alpaca, we hope to foster a deeper understanding of instruction-following models while providing researchers with a more attainable resource for their studies and explorations. This initiative marks a significant stride toward improving the overall landscape of instruction-following technologies.
-
14
GradientJ
GradientJ
Accelerate innovation and optimize language models effortlessly today!
GradientJ provides an extensive array of tools aimed at accelerating the creation of large language model applications while also supporting their sustainable management. Users have the ability to explore and optimize their prompts by preserving various iterations and assessing them according to recognized benchmarks. Furthermore, the platform allows for the efficient orchestration of complex applications by connecting prompts and knowledge bases into advanced APIs. In addition, enhancing the accuracy of models is possible through the integration of personalized data resources, which significantly improves overall functionality. This versatile platform not only enables developers to innovate but also fosters an environment for the ongoing refinement of their models, encouraging continuous improvement in their applications. By utilizing these features, developers can stay ahead in the rapidly evolving landscape of language model technology.
-
15
PanGu Chat
Huawei
Experience seamless conversations with intuitive, human-like AI interaction.
Huawei has developed an AI chatbot called PanGu Chat, designed to engage in conversations that closely resemble human interaction and respond to questions in a way akin to ChatGPT. This innovative technology seeks to improve user experience by mimicking the flow of natural dialogue, making interactions more intuitive and relatable. As a result, users can expect a more seamless communication experience when utilizing this advanced tool.
-
16
LTM-1
Magic AI
Revolutionizing coding assistance with unparalleled context and accuracy.
Magic’s innovative LTM-1 technology enables context windows that are 50 times greater than the standard ones found in traditional transformer models. Consequently, Magic has created a Large Language Model (LLM) capable of efficiently handling extensive contextual information for generating recommendations. This breakthrough empowers our coding assistant to thoroughly examine and utilize your entire code repository. By drawing on a wealth of factual knowledge and its own previous interactions, larger context windows greatly improve the accuracy and cohesiveness of AI-generated responses. We are enthusiastic about the possibilities this research presents for enhancing user experiences in coding assistance tools, paving the way for smarter, more intuitive interactions. Ultimately, we believe these advancements will significantly transform how developers engage with their coding environments.
-
17
Reka
Reka
Empowering innovation with customized, secure multimodal assistance.
Our sophisticated multimodal assistant has been thoughtfully designed with an emphasis on privacy, security, and operational efficiency. Yasa is equipped to analyze a range of content types, such as text, images, videos, and tables, with ambitions to broaden its capabilities in the future. It serves as a valuable resource for generating ideas for creative endeavors, addressing basic inquiries, and extracting meaningful insights from your proprietary data. With only a few simple commands, you can create, train, compress, or implement it on your own infrastructure. Our unique algorithms allow for customization of the model to suit your individual data and needs. We employ cutting-edge methods that include retrieval, fine-tuning, self-supervised instruction tuning, and reinforcement learning to enhance our model, ensuring it aligns effectively with your specific operational demands. This approach not only improves user satisfaction but also fosters productivity and innovation in a rapidly evolving landscape. As we continue to refine our technology, we remain committed to providing solutions that empower users to achieve their goals.
-
18
Samsung Gauss
Samsung
Revolutionizing creativity and communication through advanced AI intelligence.
Samsung Gauss is a groundbreaking AI model developed by Samsung Electronics, intended to function as a large language model trained on a vast selection of text and code. This sophisticated model possesses the ability to generate coherent text, translate multiple languages, create a variety of artistic works, and offer informative answers to a broad spectrum of questions.
While Samsung Gauss is still undergoing enhancements, it has already proven its skill in numerous tasks, including:
Adhering to directives and satisfying requests with thoughtful attention.
Providing comprehensive and insightful answers to inquiries, no matter how intricate or unique they may be.
Generating an array of creative outputs, such as poems, programming code, scripts, musical pieces, emails, and letters.
For example, Samsung Gauss is capable of translating text between many languages, including English, French, German, Spanish, Chinese, Japanese, and Korean, and can also produce functional code tailored to specific programming requirements. Moreover, as its development progresses, the potential uses of Samsung Gauss are expected to grow extensively, promising exciting new possibilities for users in various fields.
-
19
Flip AI
Flip AI
Revolutionize incident response with unmatched observability and efficiency.
Our cutting-edge model possesses the ability to understand and evaluate all types of observability data, including unstructured content, which allows for a rapid restoration of the health of software and systems. It has been meticulously crafted to effectively manage and resolve various critical incidents across multiple architectural frameworks, offering enterprise developers access to unmatched debugging capabilities. This model specifically addresses one of the most daunting challenges in software engineering: troubleshooting issues that occur in production environments. It operates efficiently without any prior training and is compatible with any observability data platform. Furthermore, it can evolve based on user input and improve its strategies by learning from past incidents and patterns unique to your setup, all while safeguarding your data. As a result, this empowers you to address critical incidents using Flip in just seconds, thereby optimizing your response time and enhancing operational efficiency. With these advanced features, you can greatly improve the resilience and reliability of your systems, ensuring a more robust infrastructure for your organization. Ultimately, this model represents a significant leap forward in the realm of software debugging and incident response.
-
20
VideoPoet
Google
Transform your creativity with effortless video generation magic.
VideoPoet is a groundbreaking modeling approach that enables any autoregressive language model or large language model (LLM) to function as a powerful video generator. This technique consists of several simple components. An autoregressive language model is trained to understand various modalities—including video, image, audio, and text—allowing it to predict the next video or audio token in a given sequence. The training structure for the LLM includes diverse multimodal generative learning objectives, which encompass tasks like text-to-video, text-to-image, image-to-video, video frame continuation, inpainting and outpainting of videos, video stylization, and video-to-audio conversion. Moreover, these tasks can be integrated to improve the model's zero-shot capabilities. This clear and effective methodology illustrates that language models can not only generate but also edit videos while maintaining impressive temporal coherence, highlighting their potential for sophisticated multimedia applications. Consequently, VideoPoet paves the way for a plethora of new opportunities in creative expression and automated content development, expanding the boundaries of how we produce and interact with digital media.
-
21
Aya
Cohere AI
Empowering global communication through extensive multilingual AI innovation.
Aya stands as a pioneering open-source generative large language model that supports a remarkable 101 languages, far exceeding the offerings of other open-source alternatives. This expansive language support allows researchers to harness the powerful capabilities of LLMs for numerous languages and cultures that have frequently been neglected by dominant models in the industry.
Alongside the launch of the Aya model, we are also unveiling the largest multilingual instruction fine-tuning dataset, which contains 513 million entries spanning 114 languages. This extensive dataset is enriched with distinctive annotations from native and fluent speakers around the globe, ensuring that AI technology can address the needs of a diverse international community that has often encountered obstacles to access. Therefore, Aya not only broadens the horizons of multilingual AI but also fosters inclusivity among various linguistic groups, paving the way for future advancements in the field. By creating an environment where linguistic diversity is celebrated, Aya stands to inspire further innovations that can bridge gaps in communication and understanding.
-
22
Tune AI
NimbleBox
Unlock limitless opportunities with secure, cutting-edge AI solutions.
Leverage the power of specialized models to achieve a competitive advantage in your industry. By utilizing our cutting-edge enterprise Gen AI framework, you can move beyond traditional constraints and assign routine tasks to powerful assistants instantly – the opportunities are limitless. Furthermore, for organizations that emphasize data security, you can tailor and deploy generative AI solutions in your private cloud environment, guaranteeing safety and confidentiality throughout the entire process. This approach not only enhances efficiency but also fosters a culture of innovation and trust within your organization.
-
23
Command R
Cohere AI
Enhance productivity and accuracy with advanced AI document insights.
Command's model generates outputs that include accurate citations, which significantly minimize the potential for misinformation while offering additional context from the original materials. It excels in various tasks such as crafting product descriptions, aiding in email writing, and suggesting sample press releases, among other functions. Users can interact with Command by posing multiple questions about a document to categorize it, extract specific details, or tackle general inquiries regarding the content. Addressing several questions related to a single document not only conserves valuable time but also applying this method to thousands of documents can result in considerable time savings for businesses. This collection of scalable models strikes an impressive balance between exceptional efficiency and solid accuracy, enabling organizations to evolve from initial experimentation to fully functional AI applications. By harnessing these advanced capabilities, companies can effectively boost their productivity and refine their operational workflows. In today's fast-paced business environment, such tools are indispensable for maintaining a competitive edge.
-
24
CodeGemma
Google
Empower your coding with adaptable, efficient, and innovative solutions.
CodeGemma is an impressive collection of efficient and adaptable models that can handle a variety of coding tasks, such as middle code completion, code generation, natural language processing, mathematical reasoning, and instruction following. It includes three unique model variants: a 7B pre-trained model intended for code completion and generation using existing code snippets, a fine-tuned 7B version for converting natural language queries into code while following instructions, and a high-performing 2B pre-trained model that completes code at speeds up to twice as fast as its counterparts. Whether you are filling in lines, creating functions, or assembling complete code segments, CodeGemma is designed to assist you in any environment, whether local or utilizing Google Cloud services. With its training grounded in a vast dataset of 500 billion tokens, primarily in English and taken from web sources, mathematics, and programming languages, CodeGemma not only improves the syntactical precision of the code it generates but also guarantees its semantic accuracy, resulting in fewer errors and a more efficient debugging process. Beyond just functionality, this powerful tool consistently adapts and improves, making coding more accessible and streamlined for developers across the globe, thereby fostering a more innovative programming landscape. As the technology advances, users can expect even more enhancements in terms of speed and accuracy.
-
25
Gen-3
Runway
Revolutionizing creativity with advanced multimodal training capabilities.
Gen-3 Alpha is the first release in a groundbreaking series of models created by Runway, utilizing a sophisticated infrastructure designed for comprehensive multimodal training. This model marks a notable advancement in fidelity, consistency, and motion capabilities when compared to its predecessor, Gen-2, and lays the foundation for the development of General World Models.
With its training on both videos and images, Gen-3 Alpha is set to enhance Runway's suite of tools such as Text to Video, Image to Video, and Text to Image, while also improving existing features like Motion Brush, Advanced Camera Controls, and Director Mode. Additionally, it will offer innovative functionalities that enable more accurate adjustments of structure, style, and motion, thereby granting users even greater creative possibilities. This evolution in technology not only signifies a major step forward for Runway but also enriches the user experience significantly.