sequenced.ai
Articles/Models & infrastructure/Blueprint//8 min read

Thinking Machines Lab pairs Inkling models with Tinker training

Understand Inkling’s open weights and Tinker’s training API, usage pricing, checkpoint portability and production-inference limits.

By Sequenced deskAI-assisted, source-led · how we work
Visit Thinking Machines Lab website ↗
InklingOpen modelsText, image and audio input.
TinkerTraining APICustomer algorithms on managed compute.
LoRAAdaptation methodTrain adapters rather than all model weights.
CheckpointsPortable outputDownload adapters for another serving environment.
Thinking Machines Lab mark
Thinking Machines Labthinkingmachines.ai · independent research

Represent this company? Verify your work email to access its workspace, or send the desk a factual correction.

Thinking Machines Lab offers its Inkling family of open-weight models and Tinker, a managed API for adapting models. The combination matters for teams that need behavior tailored to a task rather than another general-purpose chat interface. Inkling supplies a model family; Tinker provides training primitives, managed computation and checkpoint handling for Inkling and other supported open models. The developer still owns the data, learning objective and evaluation that determine whether customization is useful.

In brief
  1. 01Best fit Researchers and engineering teams with a measurable post-training problem.
  2. 02Implementation focus Keep the experiment’s data and algorithm under developer control while outsourcing distributed computation.
  3. 03Availability caveat Tinker’s serverless inference is separately labelled beta and is not recommended for intensive production use.

01 / ProductOpen models alongside programmable post-training

The Inkling overview presents Inkling and Inkling-Small as models accepting text, images and audio. Both use mixture-of-experts architectures. The page distinguishes their advertised one-million-token model context from the 64K or 256K configurations available through Tinker. Those are different deployment limits, so an application should use the limit of its chosen route rather than the largest number in a model announcement.

The original Inkling release describes a model trained from scratch with downloadable weights, and the Inkling-Small release establishes a smaller member of the same family. The company’s performance comparisons are vendor evaluations. They provide candidate-selection context, not proof that either model will handle a customer’s documents, audio or tool workflows reliably.

The Tinker documentation describes a deliberately low-level abstraction. A developer writes a Python loop, data preparation and learning logic; Tinker handles distributed GPU execution. Its current adaptation method is low-rank adaptation, or LoRA, rather than updating every base-model parameter. The distinction affects what is trained, what is downloaded and how the result is later served.

02 / AudienceFor a team that knows what better behavior means

Tinker fits a team able to specify a concrete change in model behavior. Examples include extracting fields in a consistent format, following a specialist annotation convention or learning a tool-selection policy from reviewed examples. Such projects need representative data and an evaluation set that remains separate from training. Without those, a convenient API can make it easier to produce an unmeasured model change.

It is less suitable as a first response to missing or outdated knowledge. If the task needs the latest internal document, retrieval may solve the problem without training. If instructions already produce acceptable outputs, improved prompts or validation may be cheaper to maintain. The useful question is whether repeated failures reflect behavior that examples can teach, rather than information the model never received.

Our Together AI blueprint discusses a broader open-model platform and its adaptation and serving routes. Our Fireworks AI blueprint provides another infrastructure perspective. Compare control over training logic, supported base models and the eventual serving arrangement; the easiest training experiment is not automatically the simplest production system.

03 / WorkflowProposed experiment for consistent support-ticket extraction

This proposed workflow has not been performed by Sequenced. Start with sanitized support tickets and a target schema: product, issue category, urgency and evidence span. Include ambiguous tickets, multiple issues in one message and cases that should remain unclassified. Ask a human reviewer to resolve the annotation standard before training a model to imitate it.

Establish a baseline from the chosen base model using the same prompts and output validation intended for the adapted version. Hold back examples from distinct customers or time periods where practical, so near-duplicate tickets do not leak across the split. Measure missing fields, unsupported urgency assignments and invalid output separately; a single average obscures different kinds of failure.

Use the current quickstart to establish SDK access and billing, then create a training client for a currently supported model. The guide separates service, training, sampling and administrative clients. Data preparation remains in the developer’s process, but tokens and loss inputs sent for computation reach the service. Local preparation is not equivalent to keeping all training content off the provider’s systems.

The supervised-learning cookbook documents dataset builders, model selection, learning configuration, evaluation cadence and checkpoint saving. For the pilot, train on the reviewed target responses and evaluate at fixed checkpoints. Select a checkpoint using held-out behavior rather than the smallest training loss; memorizing examples is not the same as handling new tickets.

Compare the adapted model with the baseline on the untouched set. Inspect cases where it became more confident but less faithful to the ticket. An extraction system should preserve uncertainty, especially when urgency influences staff workload. Add a regression set for refusals, malformed input and unexpected languages so a narrow improvement does not quietly damage useful general behavior.

Finally, download a sampler checkpoint and prove that the intended serving environment can load it. The checkpoint guide explains that this archive contains a LoRA adapter, not a complete standalone base model. Keep the base-model identifier, adapter configuration, tokenizer assumptions and evaluation report together. Portability needs a working deployment rehearsal, not merely a downloaded archive.

04 / PricingSeparate training, sampling and stored checkpoints

Tinker’s rate card lists prices per million tokens; the Tinker FAQ confirms US dollars. The rate card separates charges for prefill, sampling and training. Rates depend on the model and sometimes its context configuration. Cached prefill has its own discounted rate. Training-client forward-only passes are billed at the training rate, a relevant detail when an experiment performs repeated scoring without a parameter update.

The table below uses Qwen3-8B as a concrete, currently listed training example. It is a budgeting reference rather than a recommendation to choose that model. Other entries carry temporary discounts or announced retirement dates; copy the exact model identifier and current rate when approving an experiment. Do not assume a historical cookbook example remains available indefinitely.

Checkpoint storage is separate from token work. Set retention intentionally and delete experimental states that no longer need to be resumed. A small training run can leave several adapters and optimizer states behind, while repeated evaluation adds sampling cost. For a fair pilot budget, include the baseline, unsuccessful runs, held-out evaluation and the final serving rehearsal.

ItemPublished basisBudget implication
Qwen3-8B prefill$0.195; cached $0.039Input for sampling; cache eligibility matters.
Qwen3-8B sampling$0.60Generated output is a separate charge.
Qwen3-8B training$0.44Training-client forward-only passes use this rate too.
Checkpoint storage$0.10 per GB-monthSet TTL or delete unwanted checkpoints.
Serverless inferenceBeta; Inkling family onlyDo not equate experimental access with production suitability.

Selected rates from Tinker Models & Pricing, consulted 11 October 2026. USD per million tokens unless otherwise shown; Qwen3-8B is an example, not a universal rate.

05 / DistinctionsControl stays with the experiment designer

Tinker’s distinctive contribution is the boundary it draws around distributed computation. The developer can express a learning procedure without first building a GPU scheduler. That can make an unusual loss function or reinforcement-learning experiment more approachable, while leaving the difficult scientific decisions visible. A managed training API does not decide whether the reward measures the behavior the organization actually wants.

The checkpoint model also separates continuing a training run from sampling the trained result. A resumable state includes information needed by the optimizer; a sampler checkpoint is designed for inference and export. Choosing the right artifact avoids an expensive surprise when someone tries to resume from a file that only supports serving.

Inkling gives the lab a first-party model family within that workflow, but Tinker is not restricted to Inkling. Evaluate a small, economical baseline before selecting a much larger model. If the task is a narrow extraction convention, the model’s general benchmark position may matter less than its measured compliance with the schema and ability to abstain.

06 / QuestionsProduction inference needs its own decision

The current pricing documentation explicitly labels serverless inference beta, limits that offering to Inkling and Inkling-Small, and advises against intensive production use until it leaves beta. That warning is distinct from whether a model can be trained or sampled. A production application should establish a supported hosting route, capacity and service expectations independently.

Confirm the relevant base-model licence before moving an adapter outside Tinker. Access to a training endpoint does not erase the base model’s conditions or settle rights in the customer’s dataset. Also test how prompts are rendered: a model expecting a chat template can respond differently when given raw continuation text, even though both requests are syntactically valid.

For evaluation, retain examples of failures as well as headline results. Human annotation quality, a changing ticket mix and output-parser behavior can each dominate the eventual outcome. The platform makes experimentation practical; the organization still needs a process for recognizing when a previously useful adaptation should be retrained, replaced or rolled back.

07 / DecisionBegin with a measurable adaptation problem

Thinking Machines Lab is most useful to evaluate through a complete path from a defined behavior gap to a reproducible adapted model. Choose the base model, establish a baseline, train with a bounded budget and prove deployment of the selected artifact. That creates evidence for further investment without confusing training access, open weights and production inference availability.

01

Have labeled examples of a repeated failure

Build a baseline and held-out set, then run a small LoRA experiment with an explicit spending limit.

Test adaptation
02

Need control over a novel learning algorithm

Assess the API primitives and cookbook against the loss function, sampling pattern and evaluation loop you need.

Evaluate training control
03

Need dependable production inference now

Choose a confirmed serving arrangement and prove adapter compatibility; do not rely on the separately labelled beta inference offer.

Resolve serving first
What should we explore next?

A business worth understanding.

Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.

Suggestions are free. Selection and publication stay with the desk.

Sources
Filed under Models & infrastructureCompany Thinking Machines LabNot affiliated with Thinking Machines LabRequest a correctionRequest a refresh by email

Continue reading

All in this category