STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
Structured visual exploration for reinforcement-learning post-training of large multimodal models on video question answering.
Researcher-builder working on foundation models, post-training, multimodal learning, video understanding, and efficient model design.
A small selection of public projects. Project cards intentionally link only to publicly available information.
Structured visual exploration for reinforcement-learning post-training of large multimodal models on video question answering.
Stacked temporal attention inside the vision encoder to improve temporal understanding and reasoning in Video-LLMs.
A post-training compression method that shares weights across Transformer blocks and models their differences with low-rank deltas.
Adaptive token sampling for reducing redundant computation in Vision Transformers.
Automatically refreshable at deployment time.
Last value in source: 2026-09-14
German grading scale: 1.0 is the highest regular grade and 4.0 is the minimum grade. A grade of 0.0 denotes the highest doctoral distinction.