Thinking Machines Lab
Thinking Machines Lab's debut — an open-weights multimodal generalist with a million-token context.
1.048576M tokens
Konteks
131,072 tokens
Output maks
Sedang
Kecepatan
A book, an archive or a codebase in one thread.
Images and text in the same conversation, natively.
Pulling one answer out of many long sources.
A capable generalist rather than a narrow specialist.
Tuned to this model — click any line to copy.
Inkling is Thinking Machines Lab's first public model, released on 15 July 2026, and it arrived as an open-weights release rather than a closed API — which is a statement of intent from a lab founded by people who built a lot of what the field runs on.
The architecture is a mixture-of-experts transformer with about 975 billion total parameters and 41 billion active per token, and a context window of up to a million tokens. It is natively multimodal rather than multimodal by adapter: it was pretrained on roughly 45 trillion tokens spanning text, images, audio and video, and images and video frames enter the same decoder as text rather than being translated into it. There are some genuinely unusual design choices inside, including short convolutions in every decoder block and a learned relative-position bias instead of the rotary embeddings almost everything else uses.
What that adds up to in practice is a generalist with unusual reach. A million tokens is enough for a book, a long research archive or an entire codebase in one conversation, and because the multimodality is native you can put a diagram or a recording into the same thread and ask about it directly. The lab also exposed controllable thinking effort, so the model can be pushed to reason harder on the turns that deserve it.
Thinking Machines released a lighter Inkling-Small alongside it. On NinjaChat, Inkling is included in every plan — pick it from the model list and give it something large.
Dari nol ke hasil pertama Anda dalam waktu kurang dari satu menit.
01
Create a NinjaChat account and choose a plan
02
Open chat and pick Inkling from the model list
03
Paste or attach the whole source, not a summary
04
Ask it to think harder on the turns that matter
Perbandingan jujur — di mana {model} unggul, dan di mana tidak.
Native multimodal input and a larger context
Nemotron is built specifically for agent orchestration
Open weights with a million-token context
Kimi K3 is tuned harder for agentic coding
Multimodal by design rather than text-first
Qwen 3.8 Max is the larger sparse model
Setiap paket NinjaChat sudah termasuk 50+ model, studio gambar lengkap, dan pembuatan video.



Satu langganan mencakup semua model di NinjaChat, termasuk Inkling.