sequenced.ai
Articles/Models & infrastructure/Blueprint//8 min read

StepFun connects multimodal reasoning with developer tools and agents

StepFun offers reasoning, visual and audio models through APIs and a separate subscription route. Its current catalog includes Step 5 Preview and Step 3.7 Flash.

By Sequenced deskAI-assisted, source-led · how we work
Visit StepFun website ↗
Step 5 PreviewLong-context modelText, image and video input with text output.
Step 3.7 FlashMultimodal reasoningA separate model for visual and agent tasks.
StepAudioSpeech familySpeech services use their own capabilities and billing units.
Step PlanSubscription routeMonthly credits use a separate API base URL.
StepFun mark
StepFunstepfun.com · independent research

Represent this company? Verify your work email to access its workspace, or send the desk a factual correction.

StepFun develops AI models for reasoning, coding, visual understanding and speech. Its developer platform provides the components for an application rather than a complete business process. The useful question is which model and access route fit the job: the current Step 5 model is explicitly a preview, Step 3.7 Flash supports visual input, and the older 3.5 Flash remains text-only.

In brief
  1. 01Best fit Developers building applications that combine visual evidence, language reasoning and controlled tool calls.
  2. 02Commercial distinction Standard API usage and Step Plan subscriptions have separate endpoints and balances.
  3. 03Research scope Official documentation reviewed; the proposed workflow below has not been executed or benchmarked.

01 / ProductA model portfolio with different input and delivery boundaries

The platform overview presents reasoning and StepAudio services alongside a browser Studio. These belong to one StepFun company identity. Speech recognition, live speech interaction and text generation should still be evaluated as distinct interfaces; choosing one provider does not make their input formats or billing units interchangeable.

The Step 5 Preview specification describes text, image and video input, text output and a one-million-token context window. It also explains that search, code execution and tools come from the surrounding application. A model can propose a useful investigation without possessing access to your repository, browser or customer system.

The Step 3.7 Flash guide documents a 256K context window and three reasoning-effort settings. Its architecture is described as a sparse mixture of experts, with 198 billion total and 11 billion active parameters. Those are vendor specifications, not measurements of the memory needed by your integration or proof of superior task accuracy.

02 / AudienceVisual evidence makes the strongest starting point

A product engineering team receiving screenshots, short recordings and written bug reports is a concrete audience. It needs a reproducible issue description, the relevant application state and a plausible investigation path. Multimodal input could reduce the work of manually translating a recording into text, especially when the important clue is a transient error or a changed button label.

This is a less direct fit for a team wanting a finished ticketing system with ownership, service targets and deployment approval already designed. Those decisions live outside the model. Nor should a preview endpoint become the only route for a time-critical workflow until the team has checked its available capacity, revision policy and acceptable fallback behavior.

The Z.ai blueprint is useful when the input is primarily a document and a separate parsing stage may help. The Mistral AI blueprint offers another developer ecosystem for model and document workflows. Compare the accepted issue record and reviewer effort, rather than treating context length as the outcome.

03 / WorkflowA proposed assistant turns recordings into reviewable issues

This proposed pilot begins with a small collection of ordinary product defects. Use recordings made for the evaluation, scrub account details and attach the relevant build identifier. Preserve the original clip so a reviewer can inspect the event directly. Include a few reports where there is no reproducible bug; an assistant that creates an issue for every input would make the queue worse.

The quickstart distinguishes the input modalities of each model. Select 3.7 Flash for the initial visual route and treat Step 5 Preview as a separately evaluated candidate. Supply the user’s stated expectation, the visible sequence and the desired output fields. Ask for observed evidence and hypotheses in different fields so a suspected cause does not masquerade as something shown on screen.

Have the model produce a candidate title, reproduction steps, observed result and missing information. A recording may show that saving failed without showing why. Preserve that distinction. When a timestamp is supplied, review whether the referenced moment actually supports the description. A coherent account of an event is not enough if it points to the wrong interaction.

The tool-calling guide describes the application executing a requested function and returning its result to the model. Expose a narrow, read-only issue search so the assistant can find possible duplicates. Return the ticket identifier and relevant description, not the entire tracker. Treat returned arguments as requests to validate, including project scope and allowed search fields.

If a similar issue exists, let the assistant propose a relationship instead of merging records automatically. Two visually identical failures can have different causes across application versions. Show the reviewer both examples and the build identifiers before deciding whether to combine them. This avoids losing a regression inside an old closed issue.

The pilot should finish with a human-approved draft issue. Keep code changes and production deployments outside this first workflow. Measure missed reproduction steps, unsupported causal claims, incorrect duplicate suggestions and the minutes needed to approve a record. Compare those results with the existing manual process on the same set of reports.

Include a clip with unreadable text and one with multiple browser tabs. The correct response may be to request a clearer reproduction rather than infer the missing detail. Preserve the model identifier, prompt revision and request usage for the evaluation, while keeping sensitive visual material only as long as the team actually needs it.

04 / PricingToken billing and monthly credits require separate estimates

ModelUncached inputCached inputOutput
Step 5 Preview¥7¥0.35¥20
Step 3.7 Flash¥1.35¥0.27¥8.1
Step 3.5 Flash¥0.7¥0.14¥2.1

CNY per million tokens on the Chinese platform, consulted 23 September 2026 in the pricing and rate-limit table. These are standard API rates, not a universal international tariff.

The pricing page counts reasoning and final-answer tokens at the output rate. As illustrative arithmetic, one million uncached input tokens plus 100,000 output tokens on 3.7 Flash would cost ¥2.16. That is a token calculation, not a claim about how many recordings can be analyzed. Video processing, retries and longer reasoning can change the amount used by each accepted issue.

Step Plan instead provides a monthly credit pool. The entry Flash Mini plan is listed at ¥49 per month for 400 million credits; credits are not a token allowance. Monthly credits expire rather than roll over. Quarterly and annual purchases are prepaid commitments with credits still issued monthly. The documented Step Plan base URL differs from the ordinary API route, so a successful call alone does not prove it used subscription credits.

The same subscription guide distinguishes domestic payment methods from overseas Stripe payment. The CNY table above is therefore tied to the documented Chinese-platform route. Confirm the tariff and account route offered to your organization before comparing it with a provider quoting US dollars. Developer subscriptions, browser creation credits and application API costs should stay visible as separate budget items.

05 / DistinctionsTool access is explicit, and model selection remains consequential

The separation between interpretation and execution is useful for engineering work. A visual model can describe a screen and suggest a tool call; the application decides which repository, tracker or environment that request can reach. That provides a place to enforce project boundaries and preserve a reviewable record of how the draft issue was assembled.

The portfolio also allows modality to influence model selection. A text-only request need not include a video model just because the report originally arrived as a recording. Once a reviewer has accepted a concise reproduction, later classification might use text alone. Test whether that staged design reduces cost without discarding the detail that makes the issue reproducible.

Long context can hold more evidence, but it can also include irrelevant logs and superseded observations. Select the material that answers the question before increasing the window. In the proposed pilot, a short clip tied to the correct build is more useful than a large archive whose versions and timestamps have been mixed together.

06 / QuestionsPreview status and retiring services affect implementation choices

The current pricing documentation announces that the step-2x-large and step-image-edit-2 image-generation services will stop on 10 October 2026. That is consequential for a new build: visual understanding in 3.7 Flash is a different capability from those image-generation endpoints. This blueprint does not recommend starting a production dependency on the retiring services.

The general billing introduction and detailed model pricing use different descriptions of image token accounting. Estimate the selected model from its current detailed specification and returned usage, rather than assuming one fixed token count for every image across the portfolio. A short visual pilot should record actual usage before extrapolating a monthly budget.

Public documentation does not establish the account-specific capacity, availability guarantee or data terms your team will receive. Confirm those for the selected endpoint, especially the preview model. The source review establishes a plausible integration design; it does not demonstrate that StepFun identified a real bug, improved developer productivity or met a particular latency target.

07 / DecisionChoose the model after defining an acceptable issue record

StepFun merits an evaluation when visual input and controlled tool use are part of the application you want to build. Begin with the smallest workflow that ends in a useful, inspectable record. Select the model, regional route and billing channel deliberately, then expand only after the team can explain its observed errors and costs.

01

Your reports contain recordings

Evaluate visual understanding against known reproduction steps and retain the original evidence.

Pilot issue triage
02

You already use a coding subscription

Check that the configured endpoint consumes the intended credit pool and supports the selected model.

Separate the billing routes
03

You need production image generation

Account for the announced October retirement before choosing a new StepFun image-generation dependency.

Verify the active service
What should we explore next?

A business worth understanding.

Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.

Suggestions are free. Selection and publication stay with the desk.

Sources

Continue reading

All in this category