-
1
Qwen3.8-Max
Alibaba
Unleash productivity with advanced AI for complex tasks.
Qwen3.8-Max is a large-scale AI model from Qwen built for coding, coworking, research, long-horizon planning, and multimodal agent workflows. It is positioned as the most capable model in the Qwen family to date, with open weights announced for release after launch. The model uses a 2.4 trillion-parameter architecture with 95 billion active parameters and is available through QwenCloud. Qwen3.8-Max is designed to complete complex, open-ended goals end to end rather than only answer isolated prompts. In coding workflows, it can write and run code, create self-evolving harnesses, normalize requirements into issues, execute tasks through agents, run tests, trigger CI checks, and iterate through feedback. Its autonomous coding examples include a 10+ day project run, a research-paper reproduction and improvement loop, and a 24-hour online competition solution that beat most participating human teams. For professional work, Qwen3.8-Max is built to handle multi-step, tool-heavy workflows across compliance, design, food operations, engineering, rehabilitation, sports analytics, and quantitative research. The model also supports long-horizon decision-making, including autonomous chip-design optimization and extended e-commerce operations simulations. Its multimodal capabilities cover images, complex PDFs, long videos, visual production, interface inspection, frontend reconstruction, Blender visualization, interactive applications, and visual feedback loops. Qwen3.8-Max can be integrated through QwenCloud APIs and used with agent frameworks or coding assistants such as Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw. By combining agentic coding, reasoning controls, multimodal understanding, visual self-correction, long-context workflows, tool use, and open-weight availability, Qwen3.8-Max helps developers and organizations build autonomous AI systems that can produce dependable deliverables.
-
2
HappyHorse 1.1
Alibaba
Revolutionize your storytelling with enhanced AI video creation!
HappyHorse-1.1-T2V is a text-to-video model on QwenCloud built to generate high-quality videos from natural language prompts. The model supports video generation workflows where the input is text and the output is video. HappyHorse-1.1-T2V is designed with improved semantic understanding so it can more accurately interpret creative instructions. It also supports cinematic shot control, helping users guide the style, composition, and feel of generated scenes. Dynamic motion rendering helps the model produce smoother movement and more natural video sequences. The model is positioned to create richer details, stronger visual consistency, natural character actions, convincing scene atmosphere, and realistic physical dynamics. Developers can access HappyHorse-1.1-T2V through the QwenCloud API using the DashScope video synthesis endpoint. API requests can specify parameters such as resolution, aspect ratio, duration, and the text prompt. The model supports 480P, 720P, and 1080P video generation with per-second pricing, along with rate limits for requests, concurrency, and async queue tasks. HappyHorse-1.1-T2V is a hosted model rather than an open source model, and it can be tested through QwenCloud’s Try AI experience or integrated with an API key. By combining prompt-based video creation, cinematic control, motion quality, visual consistency, scalable API access, and configurable output settings, HappyHorse-1.1-T2V helps creators and developers turn ideas into generated video.
-
3
Qwen
Alibaba
Unlock creativity and productivity with versatile AI assistance!
Qwen is an advanced AI assistant and development platform powered by Alibaba Cloud’s cutting-edge Qwen model family, offering powerful multimodal reasoning and creativity tools for users at all skill levels. It provides a free and accessible interface through Qwen Chat, where anyone can generate images, analyze content, perform deep multi-step research, and build fully coded web pages simply by describing what they want. Using its VLo model, Qwen transforms ideas into detailed visuals and supports editing, style transfer, and complex multi-element image creation. Deep Research acts like an automated research partner, gathering information online, synthesizing insights, and generating structured reports in minutes. The Web Dev feature empowers users to create modern, ready-to-deploy websites with clean code using only natural language instructions. Qwen’s enhanced “Thinking” capabilities provide stronger logic, structured problem-solving, and real-time internet-aware analysis. Its Search tool retrieves precise results with contextual understanding, while multimodal intelligence enables Qwen to process images, audio, video, and text together for deeper comprehension. For developers, the Qwen API offers OpenAI-compatible endpoints, allowing seamless integration of Qwen’s reasoning, generation, and multimodal abilities into any application or product. This makes Qwen not only an AI assistant but also a versatile platform for builders and engineers. Across web, desktop, and mobile environments, Qwen delivers a unified, high-performance AI experience.
-
4
Qwen3.8-27B
Alibaba
Unlock powerful AI with practical, open-weight model flexibility.
Qwen3.8-27B is an open-weights 27B-class model connected to Alibaba’s Qwen3.8 release, built for developers, researchers, and AI teams that need a capable but more deployable model size. Alibaba’s Qwen3.8 launch described the broader model family as optimized for coding and cowork scenarios, including software development, document processing, data analysis, and professional workflows. Reports state that Alibaba planned to open-source Qwen3.8-Max alongside Qwen3.8-27B, expanding access for developers and researchers. Qwen3.8-27B gives builders a smaller alternative to the 2.4T-parameter Qwen3.8-Max model, which third-party coverage describes as Qwen’s first Max-scale model planned for open weights. The model is well suited for coding assistance, local development, agent testing, workflow automation, data analysis, document understanding, and private AI experimentation. QwenCloud documentation lists Qwen3.8-Max as supporting a 1M context window, thinking, function calling, built-in tools, and structured output, showing the broader Qwen3.8 generation’s focus on advanced agent and application workflows. Qwen3.8-27B is especially useful for teams that want Qwen-family capabilities without the infrastructure demands of Max-scale deployment. Community posts around the release point to active interest in Hugging Face, Unsloth GGUF, Ollama, and local inference use cases. Third-party coverage also notes practical hardware discussions around quantized Qwen3.8-27B deployment, including claims that 4-bit variants can fit more easily on consumer or workstation GPUs. The model can be positioned for organizations that need open AI infrastructure, coding agents, local model evaluation, private deployments, and cost-controlled experimentation. By combining open-weight access, a practical 27B model size, Qwen3.8-era performance ambitions, coding-oriented workflows, and local deployment interest, Qwen3.8-27B gives developers a flexible foundation for building AI products and agents.
-
5
Qwen-7B
Alibaba
Powerful AI model for unmatched adaptability and efficiency.
Qwen-7B represents the seventh iteration in Alibaba Cloud's Qwen language model lineup, also referred to as Tongyi Qianwen, featuring 7 billion parameters. This advanced language model employs a Transformer architecture and has undergone pretraining on a vast array of data, including web content, literature, programming code, and more. In addition, we have launched Qwen-7B-Chat, an AI assistant that enhances the pretrained Qwen-7B model by integrating sophisticated alignment techniques. The Qwen-7B series includes several remarkable attributes:
Its training was conducted on a premium dataset encompassing over 2.2 trillion tokens collected from a custom assembly of high-quality texts and codes across diverse fields, covering both general and specialized areas of knowledge. Moreover, the model excels in performance, outshining similarly-sized competitors on various benchmark datasets that evaluate skills in natural language comprehension, mathematical reasoning, and programming challenges. This establishes Qwen-7B as a prominent contender in the AI language model landscape. In summary, its intricate training regimen and solid architecture contribute significantly to its outstanding adaptability and efficiency in a wide range of applications.
-
6
Qwen2.5
Alibaba
Revolutionizing AI with precision, creativity, and personalized solutions.
Qwen2.5 is an advanced multimodal AI system designed to provide highly accurate and context-aware responses across a wide range of applications. This iteration builds on previous models by integrating sophisticated natural language understanding with enhanced reasoning capabilities, creativity, and the ability to handle various forms of media. With its adeptness in analyzing and generating text, interpreting visual information, and managing complex datasets, Qwen2.5 delivers timely and precise solutions. Its architecture emphasizes flexibility, making it particularly effective in personalized assistance, thorough data analysis, creative content generation, and academic research, thus becoming an essential tool for both experts and everyday users. Additionally, the model is developed with a commitment to user engagement, prioritizing transparency, efficiency, and ethical AI practices, ultimately fostering a rewarding experience for those who utilize it. As technology continues to evolve, the ongoing refinement of Qwen2.5 ensures that it remains at the forefront of AI innovation.
-
7
CodeQwen
Alibaba
Empower your coding with seamless, intelligent generation capabilities.
CodeQwen acts as the programming equivalent of Qwen, a collection of large language models developed by the Qwen team at Alibaba Cloud. This model, which is based on a transformer architecture that operates purely as a decoder, has been rigorously pre-trained on an extensive dataset of code. It is known for its strong capabilities in code generation and has achieved remarkable results on various benchmarking assessments. CodeQwen can understand and generate long contexts of up to 64,000 tokens and supports 92 programming languages, excelling in tasks such as text-to-SQL queries and debugging operations. Interacting with CodeQwen is uncomplicated; users can start a dialogue with just a few lines of code leveraging transformers. The interaction is rooted in creating the tokenizer and model using pre-existing methods, utilizing the generate function to foster communication through the chat template specified by the tokenizer. Adhering to our established guidelines, we adopt the ChatML template specifically designed for chat models. This model efficiently completes code snippets according to the prompts it receives, providing responses that require no additional formatting changes, thereby significantly enhancing the user experience. The smooth integration of these components highlights the adaptability and effectiveness of CodeQwen in addressing a wide range of programming challenges, making it an invaluable tool for developers.
-
8
Qwen2-VL
Alibaba
Revolutionizing vision-language understanding for advanced global applications.
Qwen2-VL stands as the latest and most sophisticated version of vision-language models in the Qwen lineup, enhancing the groundwork laid by Qwen-VL. This upgraded model demonstrates exceptional abilities, including:
Delivering top-tier performance in understanding images of various resolutions and aspect ratios, with Qwen2-VL particularly shining in visual comprehension challenges such as MathVista, DocVQA, RealWorldQA, and MTVQA, among others.
Handling videos longer than 20 minutes, which allows for high-quality video question answering, engaging conversations, and innovative content generation.
Operating as an intelligent agent that can control devices such as smartphones and robots, Qwen2-VL employs its advanced reasoning abilities and decision-making capabilities to execute automated tasks triggered by visual elements and written instructions.
Offering multilingual capabilities to serve a worldwide audience, Qwen2-VL is now adept at interpreting text in several languages present in images, broadening its usability and accessibility for users from diverse linguistic backgrounds. Furthermore, this extensive functionality positions Qwen2-VL as an adaptable resource for a wide array of applications across various sectors.
-
9
Qwen2.5-Max
Alibaba
Revolutionary AI model unlocking new pathways for innovation.
Qwen2.5-Max is a cutting-edge Mixture-of-Experts (MoE) model developed by the Qwen team, trained on a vast dataset of over 20 trillion tokens and improved through techniques such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF). It outperforms models like DeepSeek V3 in various evaluations, excelling in benchmarks such as Arena-Hard, LiveBench, LiveCodeBench, and GPQA-Diamond, and also achieving impressive results in tests like MMLU-Pro. Users can access this model via an API on Alibaba Cloud, which facilitates easy integration into various applications, and they can also engage with it directly on Qwen Chat for a more interactive experience. Furthermore, Qwen2.5-Max's advanced features and high performance mark a remarkable step forward in the evolution of AI technology. It not only enhances productivity but also opens new avenues for innovation in the field.
-
10
Qwen2.5-VL
Alibaba
Next-level visual assistant transforming interaction with data.
The Qwen2.5-VL represents a significant advancement in the Qwen vision-language model series, offering substantial enhancements over the earlier version, Qwen2-VL. This sophisticated model showcases remarkable skills in visual interpretation, capable of recognizing a wide variety of elements in images, including text, charts, and numerous graphical components. Acting as an interactive visual assistant, it possesses the ability to reason and adeptly utilize tools, making it ideal for applications that require interaction on both computers and mobile devices. Additionally, Qwen2.5-VL excels in analyzing lengthy videos, being able to pinpoint relevant segments within those that exceed one hour in duration. It also specializes in precisely identifying objects in images, providing bounding boxes or point annotations, and generates well-organized JSON outputs detailing coordinates and attributes. The model is designed to output structured data for various document types, such as scanned invoices, forms, and tables, which proves especially beneficial for sectors like finance and commerce. Available in both base and instruct configurations across 3B, 7B, and 72B models, Qwen2.5-VL is accessible on platforms like Hugging Face and ModelScope, broadening its availability for developers and researchers. Furthermore, this model not only enhances the realm of vision-language processing but also establishes a new benchmark for future innovations in this area, paving the way for even more sophisticated applications.
-
11
Qwen3
Alibaba
Unleashing groundbreaking AI with unparalleled global language support.
Qwen3, the latest large language model from the Qwen family, introduces a new level of flexibility and power for developers and researchers. With models ranging from the high-performance Qwen3-235B-A22B to the smaller Qwen3-4B, Qwen3 is engineered to excel across a variety of tasks, including coding, math, and natural language processing. The unique hybrid thinking modes allow users to switch between deep reasoning for complex tasks and fast, efficient responses for simpler ones. Additionally, Qwen3 supports 119 languages, making it ideal for global applications. The model has been trained on an unprecedented 36 trillion tokens and leverages cutting-edge reinforcement learning techniques to continually improve its capabilities. Available on multiple platforms, including Hugging Face and ModelScope, Qwen3 is an essential tool for those seeking advanced AI-powered solutions for their projects.
-
12
Qwen3-Coder
Qwen
Revolutionizing code generation with advanced AI-driven capabilities.
Qwen3-Coder is a multifaceted coding model available in different sizes, prominently showcasing the 480B-parameter Mixture-of-Experts variant with 35B active parameters, which adeptly manages 256K-token contexts that can be scaled up to 1 million tokens. It demonstrates remarkable performance comparable to Claude Sonnet 4, having been pre-trained on a staggering 7.5 trillion tokens, with 70% of that data comprising code, and it employs synthetic data fine-tuned through Qwen2.5-Coder to bolster both coding proficiency and overall effectiveness. Additionally, the model utilizes advanced post-training techniques that incorporate substantial, execution-guided reinforcement learning, enabling it to generate a wide array of test cases across 20,000 parallel environments, thus excelling in multi-turn software engineering tasks like SWE-Bench Verified without requiring test-time scaling. Beyond the model itself, the open-source Qwen Code CLI, inspired by Gemini Code, equips users to implement Qwen3-Coder within dynamic workflows by utilizing customized prompts and function calling protocols while ensuring seamless integration with Node.js, OpenAI SDKs, and environment variables. This robust ecosystem not only aids developers in enhancing their coding projects efficiently but also fosters innovation by providing tools that adapt to various programming needs. Ultimately, Qwen3-Coder stands out as a powerful resource for developers seeking to improve their software development processes.
-
13
Qwen3-Max
Alibaba
Unleash limitless potential with advanced multi-modal reasoning capabilities.
Qwen3-Max is Alibaba's state-of-the-art large language model, boasting an impressive trillion parameters designed to enhance performance in tasks that demand agency, coding, reasoning, and the management of long contexts. As a progression of the Qwen3 series, this model utilizes improved architecture, training techniques, and inference methods; it features both thinker and non-thinker modes, introduces a distinctive “thinking budget” approach, and offers the flexibility to switch modes according to the complexity of the tasks. With its capability to process extremely long inputs and manage hundreds of thousands of tokens, it also enables the invocation of tools and showcases remarkable outcomes across various benchmarks, including evaluations related to coding, multi-step reasoning, and agent assessments like Tau2-Bench. Although the initial iteration primarily focuses on following instructions within a non-thinking framework, Alibaba plans to roll out reasoning features that will empower autonomous agent functionalities in the near future. Furthermore, with its robust multilingual support and comprehensive training on trillions of tokens, Qwen3-Max is available through API interfaces that integrate well with OpenAI-style functionalities, guaranteeing extensive applicability across a range of applications. This extensive and innovative framework positions Qwen3-Max as a significant competitor in the field of advanced artificial intelligence language models, making it a pivotal tool for developers and researchers alike.
-
14
Qwen3-TTS
Alibaba
Advanced text-to-speech models for expressive, real-time voice generation.
Qwen3-TTS is a cutting-edge suite of sophisticated text-to-speech models developed by the Qwen team at Alibaba Cloud, made available under the Apache-2.0 license, which provides stable, expressive, and immediate speech synthesis, featuring capabilities such as voice cloning, voice design, and meticulous control over prosody and acoustic parameters. This collection caters to ten major languages—Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian—while also offering various dialect-specific voice profiles that allow for nuanced adjustments in tone, speech speed, and emotional expression based on the semantics of the text and the user’s directives. The design of Qwen3-TTS employs efficient tokenization and a dual-track framework, enabling ultra-low-latency streaming synthesis, with the initial audio packet produced in roughly 97 milliseconds, making it particularly suitable for interactive and real-time usage scenarios. Furthermore, the array of models provided ensures a wide range of functionalities, including quick three-second voice cloning, customization of voice qualities, and tailored voice design according to specific instructions, thereby guaranteeing adaptability for users across diverse contexts. The extensive capabilities and design flexibility of this technology underscore its potential for a multitude of applications, spanning both professional environments and personal use, paving the way for enhanced communication experiences. As such, Qwen3-TTS stands to revolutionize the way we interact with voice technologies in everyday life.
-
15
Happy Shrimp 1.0
Alibaba Cloud
Transform your ideas into fully produced songs effortlessly!
Happy Shrimp 1.0 is a groundbreaking AI-driven music creation platform that converts abstract ideas into fully formed songs from a single user prompt. This tool empowers users to kickstart the music-making process using an emotion, storytelling element, musical genre, or artistic concept, all without needing any advanced understanding of musical theory such as BPM, key signatures, or instrument selection. It is particularly skilled at generating both vocal and instrumental tracks, producing melodies, arrangements, lyrics, and vocals from scratch or using user-supplied lyrics as a foundation. With its extensive knowledge of diverse musical styles, it skillfully explores creative aesthetics from a broad spectrum of genres, cultures, and historical contexts, seamlessly transforming descriptive concepts into coherent musical compositions. The model is versatile, capable of producing music across various styles, such as Chinese music, pop, R&B, soul, hip hop, rock, funk, electronic, classical, and jazz. By tapping into its deep understanding of the world and the principles of music theory, the model articulately conveys the foundational structure and "grammar" of music, allowing it to create pieces that resonate emotionally with listeners. Ultimately, Happy Shrimp 1.0 serves as a conduit between creativity and sound, democratizing the process of music production and enabling individuals to realize their musical dreams, regardless of their prior experience. Furthermore, this innovative tool encourages creative exploration, inviting users to experiment and discover their personal sound in a supportive environment.
-
16
Qwen2.5-1M
Alibaba
Revolutionizing long context processing with lightning-fast efficiency!
The Qwen2.5-1M language model, developed by the Qwen team, is an open-source innovation designed to handle extraordinarily long context lengths of up to one million tokens. This release features two model variations: Qwen2.5-7B-Instruct-1M and Qwen2.5-14B-Instruct-1M, marking a groundbreaking milestone as the first Qwen models optimized for such extensive token context. Moreover, the team has introduced an inference framework utilizing vLLM along with sparse attention mechanisms, which significantly boosts processing speeds for inputs of 1 million tokens, achieving speed enhancements ranging from three to seven times. Accompanying this model is a comprehensive technical report that delves into the design decisions and outcomes of various ablation studies. This thorough documentation ensures that users gain a deep understanding of the models' capabilities and the technology that powers them. Additionally, the improvements in processing efficiency are expected to open new avenues for applications needing extensive context management.
-
17
Qwen3.8-Flash-Next
Alibaba
Revolutionizing AI with efficient, powerful multimodal capabilities.
Qwen3.8-Flash-Next is a pioneering open-weight multimodal Mixture-of-Experts architecture that offers an initial look at the design meant for its successor, Qwen4. This model has been expertly crafted to enhance various aspects such as attention mechanisms, residual pathways, embeddings, and optimization strategies, thereby increasing its overall functionality, enhancing computational efficiency, expanding its model capacity, and ensuring stability during training. Its unique hybrid structure combines Gated DeltaNet, which effectively condenses historical information, with Qwen Sparse Attention, facilitating the selection of meaningful context on a micro-block scale to reduce both attention and indexing expenses for lengthy sequences. The Gated Residual feature enhances the residual pathway by incorporating four streams, which helps in dynamically regulating the information flow across different layers. Moreover, the N-gram Embedding cleverly merges large-scale local-pattern memory with minimal computational overhead for each token, with the capability to transfer to host memory for added efficiency. The entire model is built around a main network comprising 125 billion parameters, supplemented by an additional 51 billion parameters specifically for N-gram embeddings, activating only 6 billion parameters for each token processed. This advanced framework underscores the continuous evolution in machine learning architectures, laying the groundwork for exciting future innovations, and it exemplifies the increasing sophistication and potential of multimodal models in various applications.
-
18
Step 5 Preview
StepFun
Empower your productivity with advanced multimodal task mastery.
Step 5 Preview epitomizes the apex of StepFun’s offerings for agentic tasks, specifically designed for practical applications in the realms of software engineering and professional knowledge, with particular excellence in financial settings. This model is adept at processing text, images, and videos, featuring an impressive 1M-token context window that suits tasks requiring extensive data, tool utilization, and continuous advancement toward specific objectives. It can effectively analyze extensive documents, amalgamate information from diverse sources, and leverage conversation threads for efficient cross-document queries and research organization. In the field of programming and software development, it showcases proficiency in a range of programming languages, assisting with debugging, code adjustments, verification tasks, and test generation. Moreover, its sophisticated multi-step agent capabilities allow applications to leverage tools for information retrieval, document analysis, detailed research, and the development of analytical reports. The model's multimodal understanding enables it to integrate images, videos, and text, facilitating tasks like chart analysis and responding to queries based on screenshots. This extensive suite of capabilities not only enhances productivity but also solidifies Step 5 Preview as an essential tool for professionals across a multitude of industries, ensuring they remain at the forefront of their respective fields.
-
19
Happy Horse
Alibaba
Transform ideas into stunning cinematic videos effortlessly!
Happy Horse is an AI video generation and editing platform designed to help creators transform prompts, images, references, and first-frame ideas into cinematic video content. The platform gives users multiple ways to begin a project, including text-based generation, reference-driven generation, first-frame input, and video editing. Creators can generate videos from imaginative concepts, then modify details to refine the final result. Happy Horse is built for visual experimentation, storytelling, and AI cinema, making it useful for artists who want to explore ideas quickly without traditional production barriers. Its creative environment includes featured projects, community videos, short AI films, and showcase content from different creators. The platform also highlights AI cinema events, encouraging users to submit and celebrate AI-made cinematic work. Users can sign in to receive free credits and take advantage of special offers for additional generation access. Happy Horse supports short-form video experimentation, concept development, visual storytelling, and creative exploration. The platform’s tools help users turn sparks of imagination into videos that can be shared, refined, or developed into larger creative projects. Its combination of generation, reference input, first-frame control, editing, and community inspiration makes it a practical workspace for AI video creators. Happy Horse helps filmmakers, designers, artists, and everyday creators bring visual ideas to life with speed, flexibility, and expressive control.
-
20
Qwen3.8-2.4T-A95B
Alibaba
Unleashing unparalleled capabilities for complex, multi-step tasks.
Qwen3.8-2.4T-A95B emerges as the largest open model in the Qwen3.8 series, presenting advanced Qwen-Max-class capabilities in a format that is accessible to the public. Built on the robust foundation of Qwen3.5, this model offers marked improvements in performance across various domains, including coding, professional applications, research, and complex, extended agentic tasks, underscoring its ability to reliably execute intricate, multi-step workflows to completion. With its innovative mixture-of-experts architecture, it features a remarkable total of 2.4 trillion parameters, of which 95 billion are activated, utilizing 512 experts and allowing for simultaneous engagement of 10 routed experts alongside one shared expert. The model supports a native context length of 262,144 tokens, extendable to about 1.01 million tokens, thereby enabling considerable adaptability for diverse applications. Additionally, enhancements in agent execution, such as superior autonomous planning and improved responsiveness to environmental cues, enhance its overall efficiency. Its extensive compatibility with popular agent frameworks and development tools further aids in smooth integration into current systems, making it an appealing option for both developers and researchers. This versatility is particularly beneficial for those seeking to leverage advanced AI capabilities in their projects.
-
21
Qwen3.8-Omni-Flash
Alibaba
Empower your productivity with advanced multimodal capabilities today!
Qwen3.8-Omni-Flash is a groundbreaking omnimodal model designed to significantly boost the efficiency of agents operating in productivity-focused settings, transitioning from basic understanding of multimodal inputs to actively performing tasks, utilizing diverse tools, and engaging in creative projects. Built upon the sophisticated Qwen3.8-Flash-Next architecture, it adeptly handles text, images, audio, and video inputs with an extraordinary context window of up to 1 million tokens, while maintaining strong performance in text-centric applications. This model transcends traditional coding and knowledge-based tasks, enriching workflows related to audio and video through capabilities such as video editing, crafting music videos, providing film commentary, summarizing audiovisual content, and facilitating real-time discussions. It particularly excels at enhancing the interpretation of long-form audio and video through organized descriptions, enabling agents to gather compelling evidence, grasp meeting content, and conduct thorough research focused on video materials. Users are empowered to specify parameters including subject matter, time frame, level of detail, and output format for video assessments, allowing for comprehensive overviews and customized analyses. This adaptability positions it as an indispensable resource for both professionals and creatives eager to optimize their productivity across a variety of multimedia platforms, ensuring that every project reaches its full potential. Furthermore, the model's seamless integration into diverse workflows opens up new possibilities for collaboration and innovation in content creation.
-
22
Qwen 4
Alibaba
Unleashing the future of AI with unparalleled intelligence.
Qwen 4 is Alibaba’s forthcoming next-generation foundation model and the planned successor to the company’s Qwen3.x model family. Alibaba announced Qwen 4 at the 2026 Apsara Conference on September 22 and confirmed that the model is currently in training. The company has not yet disclosed Qwen 4’s architecture, parameter count, context length, training-compute requirements, benchmark scores, pricing, licensing terms, or release schedule. Qwen 4 is being developed as Alibaba expands its broader AI stack across foundation models, multimodal systems, AI infrastructure, and agent-oriented cloud services. A major research direction surrounding Alibaba’s next generation of models is recursive self-improvement based on real-world tasks and empirical feedback. The company has already experimented with this approach using Qwen3.8-Max, allowing the model to participate in automated pipeline design, data validation, experimentation, error diagnosis, and post-training optimization. Alibaba reported that Qwen3.8-Max completed 33 iterative cycles during one such experiment and increased its Artificial Analysis score from 40 to 45. In a separate chip-design experiment, a Qwen model performed more than 10,000 EDA tool calls during over 60 hours of automated improvement work, illustrating Alibaba’s interest in long-horizon agentic tasks. These demonstrations describe the research program surrounding future Qwen development rather than confirmed features of Qwen 4 itself. Alibaba has also announced a longer-term roadmap in which Qwen 4.5 and Qwen 5 models are projected to reach between 5 trillion and 10 trillion parameters. Qwen 4 therefore remains a pre-release model, with detailed capabilities and access information expected to become clearer when Alibaba publishes its formal launch materials.