Here is what actually published. On 23 July 2026, application US20260212540A1 entered the public record with twenty-one claims and three independents. Claim 1 recites a system comprising an image encoder configured to transform frames from a video into visual tokens, a multimodal large language model configured to receive the visual tokens, and a temporal module interposed between the image encoder and the multimodal large language model. That is the whole of it. The assignee is NVIDIA Corp.; the named inventors are Jindong Jiang, Wonmin Byeon, Xiuyu Li and Yao Lu. It is classified under CPC G06T 9/002, G06N 3/048 and G06T 3/4046.
Published is not granted. Nothing in this document is enforceable today, the claims have not been examined on the public record, and the set that eventually issues — if one does — may look materially narrower than what is reproduced here. What follows is a description of the claims as filed, not an assessment of what they are worth.
The independents, and what separates them
Claim 1 is a system claim. Claim 8 covers the same architecture as a process: operating an image encoder to transform frames into visual tokens, and operating a multimodal large language model to receive those tokens via a temporal module interposed between the two. Claim 15 covers it a third time as a computer system comprising at least one data processor and logic configured to perform those operations. Three statutory framings of one arrangement is standard practice; it is how a drafter covers an apparatus, the method it performs, and the software that makes it perform it.
What is notable is how little claim 1 requires. It recites three components and one positional relationship. It does not specify what the temporal module computes, how the visual tokens are structured, what architecture the encoder or the language model uses, or how many frames are involved. The scope lives entirely in the arrangement — a component defined by where it sits rather than by what it does. Four separate claims use the word interposed to do that work, for the temporal module, the token compressor, the downsampling layer and the linear layer respectively.
Where the dependents get specific
The dependent claims narrow by naming things. Claims 2, 9 and 16 recite that the image encoder comprises a SigLIP model. Claims 3, 10 and 17 recite that the temporal module comprises a Mamba temporal projector. Naming externally developed, publicly known designs inside dependent claims is a deliberate choice: it produces fallback positions that are concrete and easy to compare against prior art, at the cost of being narrow enough to design around by substituting a different encoder or projector.
The system of claim 3, further comprising a token compressor interposed between the Mamba temporal projector and the multimodal large language model.— Long Video Understanding for Video-Based Visual Language Models, US20260212540A1
Claim 4, quoted above, adds the token compressor as a distinct component sitting downstream of the projector. Claim 5 then recites that the Mamba temporal projector comprises a spatial and temporal token compressor. Read together, the two place compression in two different structural positions — once as a separate stage after the projector, once as a function inside it. Claims 6 and 7 build the front end: a downsampling layer between the image encoder and the temporal module, then a linear layer between that downsampling layer and the temporal module.
The claim set carries several dependency recitations that cross statutory categories. Claim 12 is written as "The system of claim 11" while claim 11 is a process claim. Claim 13 is written as "The system of claim 8", and claim 8 is a process claim. Claim 20 is written as "The system of claim 15", where claim 15 is a computer system. These are reproduced exactly as published. They read as drafting slips of the kind routinely corrected during prosecution, and they are worth recording only because they are the sort of thing that gets tidied between publication and grant — which is precisely why the published version is not the version to rely on.
It is worth being precise about what the dependent structure implies for scope. Claims 2 and 3 both depend directly from claim 1 and can be read together, so a system using a SigLIP encoder and a Mamba temporal projector falls within both. Claim 4 depends from claim 3 rather than from claim 1, which means the token compressor limitation only ever arrives in combination with the named projector. A drafter who wanted the compressor as a standalone narrowing of claim 1 would have hung it there instead. As published, there is no intermediate claim covering an unnamed temporal module plus a separate token compressor, and that gap in the dependency chain is more consequential to the eventual scope than any of the cross-category slips above.
The application is titled Long Video Understanding for Video-Based Visual Language Models. No claim in the set recites a video duration, a frame count, a sampling rate, or any threshold separating long video from short. Length appears in the title and motivates the compression stages, but it is not a limitation. The claims as published would read on the recited arrangement operating on video of any length. Conflating a title's framing with a claim's boundary is one of the more common ways published applications get overstated, and this one invites it.
For context on the surrounding portfolio, seven other records published under an NVIDIA assignee in the same drop, including US20260212622A1 on body-model generation, US20260211755A1 on memory-fabric transfers, US20260211796A1 on kernel debugging and US20260208764A1 on lane geometry. None of them shares this one's classification. This is the only record of the eight directed at the internal arrangement of a multimodal model, and its scope, for now, is the arrangement and nothing more.
Comments
Loading comments…