AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Thinking Machines has launched Inkling, a large-scale multimodal AI model, on Hugging Face. Its open availability and scale raise questions about performance, licensing, and practical use. Key details remain unverified.

Thinking Machines has released Inkling, a 975-billion-parameter multimodal model, on Hugging Face, offering open access to developers for processing text, images, and audio. The release highlights a significant step in large-scale AI, though many technical and licensing details remain unverified.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple modalities. Its architecture incorporates 256 experts, global and sliding-window attention, and hierarchical image patching, aimed at enabling reasoning across diverse data types.

However, the release does not include independent benchmark results, safety evaluations, or licensing specifics. The model’s hardware requirements are substantial—estimated at 2 TB of VRAM for BF16 checkpoints—making it inaccessible for most consumer systems. Instead, users are expected to utilize hosted inference services or heavy quantization.

Developers can access Inkling via supported inference engines like Transformers and llama.cpp, with potential for domain-specific fine-tuning in scientific, media, or enterprise applications. For broader safety considerations, see Protecting Teens: The Case For Safe Artificial Intelligence. The model’s size and complexity limit direct deployment to specialized hardware, raising questions about practical usability and safety assessments in real-world scenarios.

At a glance
reportWhen: announced July 2026
The developmentThinking Machines released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face, emphasizing its potential for advanced reasoning across text, images, and audio.

Implications of Inkling’s Open Multimodal Approach

Inkling’s release signifies a major advancement in large-scale multimodal AI, potentially enabling more integrated reasoning across text, images, and audio. Its open availability could accelerate research and application development, especially in domains requiring complex data analysis.

However, the high hardware demands and lack of independent evaluation mean widespread adoption faces hurdles. The model’s capabilities, safety, and licensing remain uncertain, affecting how quickly it can influence AI development and deployment.

Amazon

high VRAM GPU for AI inference

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal AI Models

Prior to Inkling, multimodal models like OpenAI’s GPT-4 and Meta’s Llama 2 have demonstrated progress in processing multiple data types, but often with restricted access or limited scale. The release of Inkling on Hugging Face, with its unprecedented parameter count and open stance, marks a notable shift in the field, emphasizing scale and multimodal integration.

Training on 45 trillion tokens across text, images, and audio suggests an ambitious scope, but the lack of independent verification and detailed benchmarks leaves questions about its relative performance and safety compared to existing models.

“This model is huge.”

— Hugging Face spokesperson

Amazon

multimodal AI model hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

It is not yet clear how Inkling performs across different workloads, especially in video processing or real-time applications. Independent benchmarks, safety assessments, and licensing details remain undisclosed, leaving questions about its practical deployment and safety assurances.

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Evaluations and Developer Testing of Inkling

Developers and research organizations are expected to begin testing Inkling through supported inference engines, which will clarify its speed, accuracy, and resource needs. Independent evaluations, benchmark disclosures, safety testing, and potential domain-specific fine-tuning will follow, shaping its future adoption and understanding.

Amazon

large-scale AI inference service

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large-scale, 975-billion-parameter multimodal AI model from Thinking Machines, capable of processing text, images, and audio within a unified framework.

Can Inkling process videos?

While the architecture includes image inputs with a temporal dimension, native video processing performance has not been evaluated, so its video capabilities remain unconfirmed.

Is Inkling available for personal use?

Full deployment requires substantial hardware—estimated at 2 TB of VRAM—making it impractical for typical consumer systems. Access is primarily through hosted inference services or heavy quantization.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing details, usage restrictions, or whether training data and code are publicly available.

How does Inkling compare to other multimodal models?

Without independent benchmarks or safety evaluations, it is difficult to compare Inkling’s performance directly to models like GPT-4 or Llama 2. Its large scale and open access are notable, but practical effectiveness remains to be seen.

Source: ThorstenMeyerAI.com

You May Also Like

OpenAI Reduces Codex Model Context Size From 372K To 272K

OpenAI has reduced the context window of its Codex model from 372,000 tokens to 272,000 tokens, impacting code generation capabilities.

Migrating A Production AI Agent To GPT-5.6: 2.2X Faster, 27% Cheaper

Migrating a production AI agent to GPT-5.6 improves performance by over twice as fast and reduces costs by more than a quarter, confirmed by the developers.

Inside The Editing Room: AI’s Role In ‘Kanton Alpin Verkehrsbetriebe’

Inside the editing room of Swiss-inspired transit project, AI tools are used for real-time synchronization and design, showcasing innovative digital craftsmanship.

The NVIDIA Earnings Preview: What Q1 FY27 Will Reveal About the AI Cycle

NVIDIA reports Q1 FY27 earnings on May 20, 2026. Key focus on revenue, demand signals, and AI infrastructure health amid a trillion-dollar order backlog.