The Annual Catalan Meeting on Computer Vision (ACMCV) brings together the Computer Vision community of Catalonia, connecting research, talent generation and industry in one day. This meeting aims to strengthen the links among the Catalan Computer Vision actors, to disseminate within the community the most relevant works that have already been published abroad and to allow students from the Master’s Degree in Computer Vision to meet with members of the Catalan computer vision community and prospective employers.
Dates & Venue
Submission deadline | September 9, 2025
Registration deadline | September 16, 2025
ACMCV 2025 | September 16, 2025
Computer Vision Center & UAB School of Engineering
Program
| Start | End | Session | Location | More information |
|---|---|---|---|---|
| 14:30 | 14:45 | Track1: Msc Defences | EE Rooms / CVC Rooms | More info |
| 14:45 | 15:30 | Track2: Msc Defences | EE Rooms / CVC Rooms | More info |
| 15:30 | 16:15 | Track3: Msc Defences | EE Rooms / CVC Rooms | More info |
| 16:15 | 17:00 | Track4: Msc Defences | EE Rooms / CVC Rooms | More info |
| 17:00 | 17:45 | Track5: Msc Defences | EE Rooms / CVC Rooms | More info |
| 17:00 | 17:30 | Accreditation & poster setup | CVC Garden | |
| 17:30 | 18:15 | Industry pitch | CVC Garden | More info |
| 18:15 | 18:45 | Poster session and networking | CVC Garden | More info |
| 18:45 | 19:30 | Keynote talk: Dr David Vázquez | CVC Garden | More info |
| 19:30 | 20:00 | Awards & Closing | CVC Garden |
Keynote talk
Enterprise Visual Understanding with Vision-Language Models: From Documents to Intelligent Agents
Vision-Language Models (VLMs) have demonstrated remarkable progress in natural image understanding and creative generation, yet their performance often falls short on enterprise-critical tasks such as document analysis, chart reasoning, workflow automation, and user interface navigation. In this talk, will be presented recent advances in adapting multimodal foundation models to enterprise applications, with a focus on text-rich visual understanding, document intelligence, and visual content–to–code generation. Also, will be introduced datasets and benchmarks such as BigDocs, BigCharts, StarFlow, and StarVector, designed to push VLMs toward real-world enterprise use cases. It will also be discussed AlignVLM, a robust architecture that bridges visual and textual representations to achieve competitive results on challenging document benchmarks. Finally, it will be highlighted how these models enable the next generation of AI agents—systems capable of reasoning, planning, and acting—by grounding natural language instructions in complex graphical user interfaces. Together, these directions illustrate a path toward enterprise-ready multimodal AI that is accurate, reliable, and adaptable.
Dr. David Vázquez
Staff Research Scientist at Google DeepMind
Resources
Organizers
Organizing Committee
- Maria Vanrell, CVC & UAB
- Josep Lladós, CVC & UAB
- Núria Martínez, CVC
- Xavier Galvez, CVC
- Aurora García, CVC
Poster Session Chair:
- Danna Xue, CVC
