Skip to content
AlphaCorp AI
Wave of light particles flowing through faint circuit traces on a dark background
Generative AI22 min read

Multimodal Models: Architecture, workflow, use cases and development

Ignas Vaitukaitis, Founder & CEO of AlphaCorp AI

AI Agent Engineer ·

Multimodal Models: Architecture, workflow, use cases and development
On this page(20)
  1. What Multimodal Models Are and How They Differ From Single-Modality AI
  2. How Multimodal Model Architecture Works: Encoders, Fusion Layers and Shared Embedding Spaces
  3. The Multimodal Workflow From Raw Input to Generated Output
  4. Use Cases Where Multimodal Models Outperform Text-Only LLMs
  5. What Does It Actually Cost to Build or Run a Multimodal Model?
  6. How to Develop a Multimodal Application: Build, Fine-Tune or Call an API
  7. Path 1: Call a hosted multimodal API
  8. Path 2: Fine-tune an open vision-language model
  9. Path 3: Train a native multimodal model
  10. Where Multimodal Models Fail: Hallucination, Modality Bias and Alignment Gaps
  11. Which Multimodal Models Lead the Field This Year and How They Compare
  12. Frequently Asked Questions About Multimodal Models
  13. Is GPT-4o a multimodal model?
  14. What is the difference between multimodal and generative AI?
  15. Are multimodal models better than LLMs?
  16. What is the difference between early fusion and late fusion?
  17. How are multimodal models trained?
  18. What is a vision-language model?
  19. How many images can I send to a multimodal model in one request?
  20. Where to Start With Your First Multimodal Project

Multimodal models are language models that take in images, video, audio and other data alongside text, built from an encoder, a fusion layer and a transformer backbone. This guide explains the three fusion designs in use, the stage-by-stage training workflow, where these models beat text-only LLMs, what they cost per visual token, and how to decide between a hosted API, fine-tuning an open multimodal model, or training natively. As of October 09, 2026, the architecture question is still open and the evidence runs against the common assumption that bolting a vision encoder onto an LLM is the safe default.

  • Google reported 81% on MMMU-Pro and 87.6% on Video-MMMU for Gemini 3 in November 2025, both vendor-measured figures from its announcement blog.
  • The Apple and Sorbonne scaling study at ICCV 2025 trained 457 native multimodal models and found no inherent advantage for late fusion over early fusion.
  • Meta’s Llama 4 trained on more than 30 trillion tokens across text and vision, per Meta’s April 2025 launch post.
  • One high-resolution image costs up to 4,784 visual tokens on Claude 4.7 and later, at one token per 28 by 28 pixel patch, per Anthropic’s vision documentation checked October 2026.
  • GPT-4o answers speech in 232 ms at minimum and 320 ms on average, according to OpenAI’s October 2024 system card.

What Multimodal Models Are and How They Differ From Single-Modality AI

A multimodal model is a language model extended to take in, and sometimes produce, images, video, audio and other data types alongside text, where single-modality AI handles one kind of input only. The research literature calls the text-plus-vision variants vision-language models (VLMs) and the broader class multimodal large language models (MLLMs). Both terms describe the same idea: one network that reads a chart, reads the question typed under it, and answers in a sentence.

A text-only LLM can’t do that. It sees tokens from a text vocabulary and nothing else. A classic vision model (an image classifier, an OCR engine) runs one task on one input type and returns a label or a string. A speech recognizer transcribes. Each is useful. None of them can hold a picture and a paragraph in the same context and reason across both, which is the gap multimodal models close.

The modalities in play as of 2026:

  • Images: photos, screenshots, scans, charts and rendered documents.
  • Video: frame sequences with timestamps. Google’s Gemini 2.5 technical report (July 2025) describes processing up to 3 hours of video in one context.
  • Audio and speech: GPT-4o (October 2024 system card) takes speech in and produces speech out without a separate transcription step.
  • Code: Google lists it as a first-class input alongside text, images, video and audio for Gemini 3 (November 2025).
  • Actions: vision-language-action models in robotics add motor commands as an output modality.

The field has also moved from understanding-only to understand-and-generate. Through 2023, most MLLMs could describe an image but only emit text. Meta FAIR’s Chameleon (May 2024) changed the framing by producing images and text in any order from a single token stream, and a 2026 ACL Findings survey now reviews “unified” MLLMs as their own category. I’d treat the understanding-only label as a historical stage. The models worth evaluating this year do both, even if your first application only needs one direction. Hugging Face’s vision-language model explainer (2024, updated 2025) remains the clearest short primer on how the understanding side is wired.

How Multimodal Model Architecture Works: Encoders, Fusion Layers and Shared Embedding Spaces

Multimodal model architecture works by converting each input type into tokens or embeddings a transformer can read, then fusing them in one of three ways: through a small connector into a pretrained LLM, through cross-attention layers inserted into a frozen LLM, or inside a single backbone trained on every modality from the first step of pre-training.

Design familyHow fusion happensNamed examples
Encoder + connector + LLM (late fusion)Vision encoder output is projected into the LLM’s embedding space and prepended to the text tokensLLaVA (CLIP encoder, projection, Vicuna decoder), Qwen3-VL, Idefics3
Cross-attention injectionA Perceiver Resampler compresses variable visual features into a fixed set of latent tokens, which gated cross-attention layers read inside a frozen LMFlamingo
Native early fusionText and visual tokens enter one backbone from the start of pre-trainingLlama 4 (April 2025), Chameleon (May 2024), GPT-4o (2024)

Fuyu-8B is the odd one out: it drops the vision encoder entirely and feeds raw image patches through a projection layer.

The building blocks show up across all three families, in different combinations.

  • Contrastive encoders: CLIP (OpenAI, 2021) trained on 400 million image-text pairs and made zero-shot transfer practical. SigLIP 2 (Google DeepMind, February 2025) adds captioning-based pre-training, self-distillation, masked prediction and online data curation, handles native aspect ratios, and ships in four sizes from 86M to 1B parameters.
  • Connectors: the original LLaVA (2023) used a single linear projection; LLaVA-1.5 (late 2023) swapped in an MLP. Cambrian-1 (NeurIPS 2024) proposed a Spatial Vision Aggregator that cuts the visual token count on the way in.
  • Token quantization: Chameleon converts images into discrete tokens so one vocabulary covers pixels and words.
  • Positional encoding for video: Qwen3-VL (November 2025) uses Interleaved-MRoPE plus DeepStack, which pulls features from several ViT layers, and ties video frames to textual timestamps across a native 256K interleaved context.
  • Omni designs: Qwen3-Omni (September 2025) splits the job into a Thinker that reasons and a Talker that produces speech, both mixture-of-experts.

Which family wins? The honest answer is that the evidence favors early fusion at small scale and says nothing settled about frontier scale. The Apple and Sorbonne scaling study of native multimodal models (ICCV 2025) trained 457 models and found no inherent advantage for late fusion. Early fusion did better at lower parameter counts and cost less to train and serve, and adding mixture-of-experts let the network learn modality-specific weights on its own. The authors ran at limited scale, so the result is a strong hint and no more.

Apple’s earlier MM1 work (March 2024) found that the image encoder, input resolution and visual token count drive results, while connector design barely moves the needle. Cambrian-1 tested more than 15 vision encoders and still argued the connector matters for token efficiency. Both can be true.

Hugging Face’s Idefics3 authors (2024) put it plainly: the field has no consensus yet on data, architecture or training recipe.

In production the split persists. Llama 4, Chameleon and GPT-4o describe unified native designs, while the Qwen3-VL family and most open models keep a ViT feeding an LLM. My reading: spend your attention on encoder choice and resolution first, and treat the connector as the last knob to turn.

The Multimodal Workflow From Raw Input to Generated Output

The multimodal workflow runs raw pixels or audio through an encoder, projects the result into the language model’s embedding space, generates autoregressively, and emits text, speech or images, with a training pipeline that builds those stages up one at a time.

The training side goes in order:

  1. Pick or pre-train the encoder. Most teams take a contrastive checkpoint such as CLIP or SigLIP 2. Meta trained Llama 4’s MetaCLIP-based encoder against a frozen Llama so its features already suit the decoder (Meta, April 2025).
  2. Align the projector. Freeze both the encoder and the LLM and train only the small connector between them, the recipe LLaVA used in 2023 and Hugging Face’s explainer still recommends as the first stage.
  3. Instruction-tune jointly. Unfreeze the decoder and train it with the projector on captions, interleaved image-text documents and plain text. Apple’s MM1 (March 2024) reported a caption to interleaved to text ratio near 5:5:1 as best for few-shot results. The text slice matters: Qwen3-VL (November 2025) reports pure-text performance matching or beating comparable text-only backbones, which is the payoff for keeping it in the mix. This stage is where most domain fine-tuning of an LLM actually happens.
  4. Post-train. Medical models such as MedMO layer reinforcement learning with verifiable rewards on top of cross-modal pre-training and multi-task instruction tuning.
  5. Evaluate. MMMU (CVPR 2024) holds 11.5K college-level problems across 30 subjects. At its late-2023 release the best scores sat under 60 percent. Those numbers are history now: Google reported 81% on MMMU-Pro and 87.6% on Video-MMMU for Gemini 3 in November 2025, vendor-measured.

At inference the path is short. An image is cut into patches, each patch becomes one visual token (28 by 28 pixels per token on Claude, up to 4,784 tokens for a single image on the newest models as of October 2026), the encoder embeds them, the projector maps them into the LLM’s space, and they sit in context ahead of the text prompt. Then the decoder runs as usual.

Two practical details trip up almost every team on its first deployment. Images placed before the question outperform the reverse ordering, per Anthropic’s vision documentation. And compression artifacts cost accuracy that no amount of prompt wording recovers, so the JPEG quality setting someone picked to save bandwidth becomes a model-quality setting.

Visual tokens dominate the bill at high resolution and on video, which is why a TMLR 2026 survey catalogs token-compression methods for image, video and audio inputs. Latency can still be low. The GPT-4o system card (October 2024) puts audio response time at 232 ms minimum and 320 ms on average, fast enough for a conversation.

Use Cases Where Multimodal Models Outperform Text-Only LLMs

Multimodal models beat text-only LLMs wherever the answer lives in pixels, waveforms or video frames that a text pipeline would otherwise have to flatten first: charts and scanned documents, robot control, medical imaging, spoken conversation and long video. On plain text documents, an OCR front end feeding a text model still costs less and runs faster.

DomainWhat the multimodal model addsEvidence and yearWhere text-only or OCR still wins
Documents and chartsReads charts, embedded screenshots and page layout directly2026 comparisons put vision models ahead of OCR-based RAG on charts and screenshots, at a higher cost per querySimple text documents, where OCR keeps the cost and latency edge
Robotics (VLA)Plans from camera input and emits motor commandsGemini Robotics 1.5 (September 2025): state of the art on 15 embodied-reasoning benchmarksA text model cannot see the workbench
Healthcare imagingReads scans alongside clinical textRadiology dominates the FDA list of AI-enabled devices, with entries through 29 June 2026Most cleared devices are task-specific imaging AI rather than general MLLMs
Voice assistantsSpeech in, speech out, no transcript in betweenGPT-4o (2024); Qwen3-Omni (September 2025) speaks 10 languagesText chat, when tone and timing don’t matter
Long videoGrounds answers to timestamps across hours of footageGemini 2.5 Pro handles up to 3 hours of video (July 2025); Qwen3-VL (November 2025) does long-video groundingTranscript search, when the visual channel adds nothing

Documents are where most enterprise teams meet this question first. Anyone who has run an OCR-based retrieval pipeline over scanned invoices knows what happens the first time a merged-cell table or a bar chart shows up: the text extraction turns to soup, and the model answers confidently from the soup. A vision-capable model reads the page as a page. That fixes the chart problem and raises your per-query bill, which is why a production RAG pipeline often routes plain text pages through OCR and sends only the visually dense pages to the multimodal model.

Robotics is the sharpest case for the technology. Google DeepMind’s Gemini Robotics 1.5 (September 2025) splits the job in two: an Embodied Reasoning model plans, and a vision-language-action model executes. Google reports that skills transferred across three different robot bodies, ALOHA 2, Apollo and Franka, without retraining per platform. Those figures come from the vendor.

Healthcare deserves a careful read. Radiology leads the FDA’s list of AI-enabled medical devices by a wide margin, and the FDA itself says the list is incomplete. Nearly all of those entries are narrow imaging tools cleared for one task. A general multimodal model that reads a chest X-ray and a clinical note together is a different product category, and the cleared-device count says little about it.

Four headline figures from vendors. 15: the number of embodied-reasoning benchmarks where Google DeepMind's Gemini Robotics 1.5 reports state of the art (September 2025). 3: the robot platforms, ALOHA 2, Apollo and Franka, that its skills transferred across without retraining. 10: the number of languages Qwen3-Omni speaks (September 2025). 3 hours: the length of video Gemini 2.5 Pro can process in one context (July 2025). The vendors measured all of these figures themselves.
Gemini Robotics 1.5 reports state of the art on 15 embodied-reasoning benchmarks, and its skills transfer across 3 different robot bodies. Source: Google DeepMind, Alibaba Qwen and Google vendor reports, 2025.

What Does It Actually Cost to Build or Run a Multimodal Model?

Running a multimodal model costs more than running a text-only one because each image or video frame becomes hundreds or thousands of extra input tokens, and building one from scratch runs to tens of trillions of training tokens. Visual tokens are the unit that matters on both sides of that ledger.

Start with inference, because that’s the bill most teams will actually pay. Anthropic’s vision documentation, checked October 2026, spells out the arithmetic for Claude:

  • Token per patch: every 28 by 28 pixel patch of an image counts as one visual token.
  • Newer models: Claude 4.7 and later accept images up to 2,576 px on the long edge and spend up to 4,784 visual tokens on a single image.
  • Earlier models: the cap is 1,568 px and 1,568 visual tokens per image.

So one high-resolution image on a current model costs roughly three times the tokens of the same image on the previous generation. A twenty-page scanned contract sent at full resolution is a six-figure token request before you’ve typed the question. Video is worse, since each sampled frame is an image. That cost pressure is the entire reason a 2026 TMLR survey catalogs token-compression methods for image, video and audio inputs: the field is trying to shrink the token count without losing what’s in the picture.

Training is a different order of magnitude. Meta’s Llama 4 (April 2025) trained on more than 30 trillion tokens across text and vision. Even the encoder alone is expensive: CLIP (2021) needed 400 million image-text pairs. The Llama 4 post gives a token count and no compute bill, and that’s typical of the frontier labs.

One finding cuts the other way. The Apple and Sorbonne scaling study (ICCV 2025), across 457 trained models, found early-fusion models cheaper to train and deploy than late-fusion ones at lower parameter counts. If you’re building small, the “simpler” encoder-plus-LLM design may cost you more.

The build-versus-API tradeoff reduces to a comparison of two token streams. Calling an API means paying per visual token on every request, forever. Fine-tuning or training means paying for training tokens once, then owning the serving cost. The crossover point depends on your query volume and image resolution, and no published study settles it for you.

A horizontal bar chart of the maximum visual tokens per image on Claude models, where each 28 by 28 pixel patch counts as 1 visual token. Claude 4.7 and later: up to 4,784 visual tokens, with a maximum long edge of 2,576 px. Earlier Claude models: up to 1,568 visual tokens, with a maximum long edge of 1,568 px. Source: Anthropic vision documentation, checked October 2026.
A single high-resolution image can cost up to 4,784 visual tokens on Claude 4.7 and later, compared with 1,568 on earlier Claude models. Source: Anthropic vision documentation, 2026.

How to Develop a Multimodal Application: Build, Fine-Tune or Call an API

AlphaCorp AIonline
Let's talk

Curious what AI could do for your business?

No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.

View Services

Developing a multimodal application comes down to three paths: call a hosted API with images in the request, fine-tune an open vision-language model on your own data, or train a native multimodal model from scratch, and nearly every team should start with the first. The paths differ in what you control and what you pay for.

Path 1: Call a hosted multimodal API

This is a day-one task with a handful of rules. Using Anthropic’s October 2026 documentation as the reference setup:

  1. Send each image as base64, a public URL or a Files API file_id.
  2. Place images before the text prompt in the message. Image-first ordering gets better answers.
  3. Stay within the per-request limit: up to 600 images, or 100 on models with a 200k context window.
  4. Resize on your side to the model’s resolution cap, so you control what gets downsampled instead of the API.

That’s it. You’ll have a working prototype in an afternoon and an evaluation set by the end of the week.

Path 2: Fine-tune an open vision-language model

Pick this when a hosted model keeps failing on your specific visuals. The Qwen3-VL family (November 2025) gives you a size ladder: dense models at 2B, 4B, 8B and 32B parameters, plus mixture-of-experts models at 30B-A3B and 235B-A22B. The recipe hasn’t changed much since LLaVA in 2023. Freeze the encoder and the LLM, train the projector on your image-text pairs, then unfreeze the decoder for joint instruction tuning. Swapping the encoder is also on the table: SigLIP 2 (February 2025) ships four checkpoints from 86M to 1B parameters, and Cambrian-1 (NeurIPS 2024) released weights, data and training recipes you can copy outright.

Path 3: Train a native multimodal model

Reserve this for when no hosted or open model offers the modality you need, or when your query volume makes per-token API pricing the largest line in the budget. Native early fusion worked cheaply at small scale in the Apple and Sorbonne study (ICCV 2025), and mixture-of-experts let the network learn modality-specific weights without hand design. At frontier scale you’re in Llama 4 territory: 30 trillion tokens and a research team.

The AlphaCorp AI three-question rule for choosing a path:

  • Does a hosted model already pass your evaluation set? Call the API and ship.
  • Does it fail on your forms, scans or screenshots, and do you own a labeled set of them? Fine-tune a Qwen3-VL-class open model.
  • Do you need a modality nobody hosts, or does the API bill exceed a training run? Only then train natively.

Where Multimodal Models Fail: Hallucination, Modality Bias and Alignment Gaps

Multimodal models fail in four recurring ways: they describe things that aren’t in the image, they let the text prompt override what the pixels show, they stumble on degraded or tiny inputs, and they can be jailbroken through the image channel in ways a text filter never sees.

Hallucination comes first because it’s the failure you’ll hit in week one. A July 2025 survey of multimodal hallucination evaluation splits the problem in two. Faithfulness hallucinations contradict the input: the model names a fourth person in a photo of three. Factuality hallucinations contradict the world: the caption is consistent with the picture and still wrong about what the object is. A separate 2024 survey (revised 2025) centers the whole category on cross-modal inconsistency, meaning the text output and the image input disagree. That framing matters for mitigation. Your evaluation set needs image-grounded checks, since a text-only judge can’t catch a model that ignored the picture.

Vendor documentation is unusually candid here. Anthropic’s vision docs, current as of October 2026, list what Claude gets wrong:

  • Hallucinates more on low-quality, rotated or very small images, with under 200 px being the danger zone.
  • Returns approximate counts and approximate object locations, so “how many” and “where” are soft answers.
  • Cannot reliably tell whether an image was AI-generated.
  • Refuses to identify people by name.
  • Is not designed to interpret complex diagnostic scans such as CT or MRI.

That last line should stop any healthcare team cold before it stops a radiologist.

Image-based jailbreaks are the failure mode most teams forget to test. A 2024 survey of jailbreak attacks on multimodal generative models (revised since) documents prompts rendered as images, or adversarial pixels, slipping past safety filters tuned on text. If your deployment accepts user uploads, the image is an attack surface.

Benchmark saturation is the quieter problem. The original MMMU scores of 56 to 59 percent from late 2023 are long gone, and the harder MMMU-Pro now carries vendor-reported scores above 80 percent (Google, November 2025). A model that aces a public benchmark may still miss your invoice format, so a published score tells you where the field is and nothing about your task.

Which Multimodal Models Lead the Field This Year and How They Compare

As of October 2026, the multimodal models with verified public specs are Gemini 3, GPT-4o, Llama 4, the Qwen3-VL and Qwen3.5-Omni families, and Claude’s vision models, and they split cleanly into native early-fusion designs and encoder-plus-LLM designs. Gemini 3 holds the strongest vendor-reported benchmark figures among them.

Model (release)Modalities inContextArchitecture familyOpennessVerified figures
Gemini 3 (Google, Nov 2025)Text, images, video, audio, code1M tokensNot published in detailClosed81% MMMU-Pro, 87.6% Video-MMMU (vendor)
GPT-4o (OpenAI, 2024)Text, vision, audio in and outNot stated in system cardNative, end-to-end autoregressiveClosed232 ms minimum, 320 ms average audio latency
Llama 4 (Meta, Apr 2025)Text, imagesNot stated in launch postNative early fusion, MetaCLIP-based encoderOpen weightsTrained on 30T+ tokens
Qwen3-VL (Alibaba, Nov 2025)Text, images, video256K native interleavedViT encoder feeding LLM, Interleaved-MRoPEOpen weightsDense 2B to 32B, MoE 30B-A3B and 235B-A22B
Qwen3.5-Omni (Alibaba, Apr 2026)Text, audio, audio-visualNot stated in abstractThinker-Talker MoEOpen weightsSOTA on 215 audio subtasks (self-reported)
Claude 4.7+ (Anthropic, 2026)Text, imagesModel-dependentNot publishedClosed4,784 visual tokens and 2,576 px per image

Read the “verified figures” column with care. Every benchmark number here comes from the vendor that built the model. Google’s Gemini 3 scores, Alibaba’s claim that Qwen3.5-Omni beats Gemini 3.1 Pro on key audio tasks, and Qwen3-Omni’s earlier claim of open-source state of the art on 32 of 36 audio benchmarks (September 2025) are all self-measured. None has an independent replication I’d point you to.

Two gaps are worth naming so you don’t fill them from memory. OpenAI’s GPT-5.x line is widely discussed, with third-party dates of March 2026 for GPT-5.4 and April 2026 for GPT-5.5, but OpenAI’s own specs and modality lists weren’t confirmable, so GPT-4o stands in as the last fully documented OpenAI omni model. Aggregator leaderboards for MMMU-Pro from mid-2026 place GPT-5.4 Pro, Gemini 3.1 Pro and Gemini 3.5 Flash at the top. Treat those rankings as rumor until a primary source publishes them.

Architecture is the more durable comparison. GPT-4o and Llama 4 describe single backbones trained across modalities from the start. Qwen3-VL keeps the older shape of a vision transformer feeding a language model, and remains the most complete open size ladder for anyone planning to fine-tune. On the open-weights side, the Qwen3-VL technical report (November 2025) is the one document I’d read in full before choosing a base model.

Frequently Asked Questions About Multimodal Models

Is GPT-4o a multimodal model?

Yes. OpenAI’s GPT-4o system card (October 2024) describes it as an autoregressive omni model trained end-to-end across text, vision and audio, so one network handles all three without a separate speech-to-text step. Its audio responses arrive in 232 ms at minimum and 320 ms on average, which is conversational speed. Specs for OpenAI’s later GPT-5.x models haven’t been confirmed from a primary source.

What is the difference between multimodal and generative AI?

Generative AI describes what a model does: it produces new text, images or audio. Multimodal describes what a model can take in and reason over: more than one data type at once. A model can be either, both or neither. GPT-4o is both, since it accepts images and speech and generates speech back. A classic image classifier is neither.

Are multimodal models better than LLMs?

They’re better when the answer lives in a picture, a waveform or a video frame, and they cost more per query when it doesn’t. A multi modal model reads a chart or a scanned form directly, which an OCR pipeline into a text LLM mangles on complex layouts. On plain text documents, OCR plus a text model keeps the cost and latency advantage. Pick by where your data actually lives.

What is the difference between early fusion and late fusion?

Early fusion feeds text and visual tokens into one backbone from the first step of pre-training, as Llama 4 (April 2025) and Chameleon (May 2024) do. Late fusion bolts a pretrained vision encoder onto a pretrained LLM through a small connector, the LLaVA pattern. The Apple and Sorbonne scaling study (ICCV 2025), covering 457 trained models, found no built-in advantage for late fusion and found early fusion cheaper at smaller parameter counts.

How are multimodal models trained?

In stages. A vision encoder such as CLIP or SigLIP 2 is pre-trained contrastively on image-text pairs, then a projector is trained to map its output into the language model’s embedding space while both ends stay frozen. Next the decoder is unfrozen for joint instruction tuning on captions, interleaved documents and plain text, with Apple’s MM1 (March 2024) reporting a roughly 5:5:1 mix as best for few-shot results. Some models add reinforcement learning afterward.

What is a vision-language model?

A vision-language model (VLM) is a multimodal model limited to two modalities: images (sometimes video) and text. It typically pairs a vision encoder with a language model so it can answer questions about a picture, describe it, or read text and charts inside it. LLaVA (2023), Idefics3 (2024) and Qwen3-VL (November 2025) are VLMs. Models that also handle audio, such as GPT-4o or Qwen3-Omni, are usually called omni models instead.

How many images can I send to a multimodal model in one request?

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor
Built for production

What could a custom AI agent take off your plate?

We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.

View Services

It depends on the provider. Anthropic’s vision documentation, current as of October 2026, allows up to 600 images per request, or 100 on models with a 200k context window, with each image capped at 4,784 visual tokens on Claude 4.7 and later. Placing images before the text prompt improves answers, and heavily compressed images degrade accuracy.

Where to Start With Your First Multimodal Project

Start your first multimodal project with one modality pair, one measurable task and a hosted API call, then earn your way toward fine-tuning only after an evaluation set tells you where the hosted model falls short. Image plus text is the pair most teams need, and a scanned form or a chart-heavy report is the task that pays back fastest.

Build the evaluation set before the prototype. Include images that match your real inputs, and seed it with the failure cases vendors already admit to: rotated scans, images under 200 px, heavy JPEG compression, and questions that ask for counts or locations. Judge answers against the picture itself. A text-only grader passes outputs that ignored the image entirely.

Then budget visual tokens. At 4,784 tokens for a single high-resolution image on Claude 4.7 and later (October 2026), resolution is a cost decision you should make deliberately, page by page, instead of letting a default setting make it.

Revisit open models when the API bill or the error rate forces the question, and the Qwen3-VL size ladder is the place to look. Ship the small version first. If you’d rather have engineers who have already done this on scanned documents sit beside your team, talk to AlphaCorp AI about your multimodal model project.

Share
Newsletter · Weekly

Stay Ahead of AI

One email per week with the AI engineering insights, agent builds, and tools that actually matter.

No spamUnsubscribe anytimeFree forever

In every issue
  1. 01One agent build, taken apart step by step
  2. 02The tools that earned a place in our stack this week
  3. 03What broke in production, and what we changed

Written by Ignas Vaitukaitis, founder of AlphaCorp AI.

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor

Ready to Ship
Your AI System?

Book a free call and let's talk about what AI can do for your business. No sales pitch, just a real conversation.