Video Infrastructure, Meta
Course title: "Optimization of video streaming - from past to future"
Abstract: We will review the history of video encoding, covering aspects of video codecs and quality metrics that have been exploited to optimize video delivery. Next, we will discuss the roots of convex optimization in video coding, connecting statistical multiplexing, rate-distortion optimization and how these led to the dynamic optimization framework. We will also discuss the complexity of video encoding (and quality metrics) and present the concept of a three-part tradeoff among bitrate-quality-complexity, as well as the important role this tradeoff plays in massive video processing systems. We will analyze software (SW) and hardware (HW) transcoding pipelines and conclude with an array of open research problems, including the use of ML/AI in video coding and video quality metrics.
Ioannis Katsavounidis (IEEE SM) received M.S., EEE and Ph.D. degrees in Electrical Engineering at the University of Southern California (USC), CA, USA. His work and research focused on Signal Processing, as part of the Signal and Image Processing Institute (SIPI), Caltech’s High Energy Physics department at the Italian National Laboratory at Gran Sasso (Laboratori Nazionali del Gran Sasso – LNGS), working as engineer for the MACRO (Monopoles, Astrophysics and Cosmic Ray Observatory) large-scale high-energy physics experiment. He was an Associate Professor at the Department of Electrical and Computer Engineering at the University of Thessaly in Volos, Greece (2008 to 2015); a Senior Research Scientist at Netflix in Los Gatos, CA, USA, (2015 to 2018). He is a Research Scientist at Meta Platforms in Menlo Park, CA, USA, supporting all video processing for the popular Facebook, Instagram and Messenger applications.
University of Campinas (Unicamp), Campinas, SP, Brazil
Course Title: "AI as a Guardian: From Proactive Health and Wellbeing to Digital Integrity"
Abstract: Artificial intelligence is a present-day force reshaping how we care for people and how we defend the integrity of information. In this talk, we explore two complementary frontiers where AI moves us from reactive to proactive postures. First, we introduce *Viva Bem*, a Hub for Artificial Intelligence on Health and Well-being dedicated to turning reactive healthcare into proactive solutions. We discuss how AI, coupled with wearable sensing, can help address anxiety, high blood pressure, Parkinson's disease, diabetes, sedentarism, and elder care, anticipating problems rather than merely responding to them. Second, we turn to *digital forensics* and its much-needed alliance with AI. Offering a panoramic view of the last ten years, we trace the field's evolution from basic computer vision and simple NLP toward powerful AI-driven methods, covering tampering detection, fact-checking, authorship analysis, and recent advances in self-supervised learning for large-scale forensic search. Across both domains, one question remains: AI is here to stay, how do we best leverage it to protect, empower, and uplift our society?
Anderson Rocha is Full-Professor of Artificial Intelligence and Digital Forensics at the Institute of Computing, University of Campinas (Unicamp), Brazil. He is the Head of the Artificial Intelligence Lab., Recod.ai, at Unicamp and was the former Director of the Institute for the 2019-2023 term. He has served as an elected member of the IEEE Information Forensics and Security Technical Committee (IFS-TC). He is a Microsoft Research and a Google Research Faculty Fellow. He is ranked among the Top 2 of research scientists worldwide, according to PlosOne/Stanford and Research.com studies. Finally, he is now a LinkedIn Top Voice in Artificial Intelligence for continuously raising awareness of Al and its potential impacts on society at large.
Technische Universität, München, Germany
Course Title: "Encoding Touch: Acquisition, Compression, and Delivery of Haptic Signals"
Abstract: Touch is a fundamental component of human interaction with the physical world and is emerging as a key sensing and communication modality in robotics, immersive media, and teleoperation. Unlike conventional multimedia signals, haptic information is embedded in closed-loop interactions between humans and their environment, imposing stringent requirements on latency, bandwidth, and system stability. This lecture presents the complete signal processing pipeline for haptic communication, covering the acquisition of kinesthetic and tactile signals, their efficient representation and perceptually motivated compression, transmission over communication networks, and high-fidelity rendering at the receiver. The talk will discuss how models of human haptic perception enable more efficient encoding of touch and will highlight recent advances in haptic codecs, as well as recent standardization efforts toward interoperable haptic communication systems.
Eckehard Steinbach is currently a Full Professor for Media Technology. His current research interests are in the area of audio-visual-haptic information processing and communication as well as networked and interactive multimedia systems. In March 2005 Prof. Steinbach has been appointed as a guest professor at the Chinesisch-Deutschen Hochschulkolleg (CDHK) at Tongji University in Shanghai. Prof. Steinbach and his team have received several best paper, best student paper or best poster awards for their work. Prof. Steinbach is the recipient of the 2011 “Forschungspreis Technische Kommunikation” of the Alcatel-Lucent Foundation. He was elected Fellow of the IEEE in 2015 for his contributions to visual and haptic communications.
University of Trento, Italy
Course Title: "Don’t Guess, Escalate: Towards Explainable Uncertainty-Calibrated AI Forensic Agents"
Abstract: The rapid evolution of AI-driven content synthesis is reshaping the landscape of multimedia forensics and consequently our ability to distinguish authentic media from synthetic content. In this lecture we will review recent advancements in the field and present a proposal for a paradigm shift toward AI forensic agents. In this vision they will act as reliable orchestrators that select and combine forensic detectors, reconstruct provenance and context, reason about when their own tools can be trusted, and calibrate uncertainty to abstain or escalate when evidence is insufficient. We will discuss with students about the necessity of prior over the reliability of detectors, how to integrate this prior with the media context, and how we could plan to effectively adapt forensic results into reports targeting different users (journalists, lawyers, general users, etc.). We review early implementations of forensic agents and identify where they fall short, define the capabilities a trustworthy agent requires, and call for golden-set benchmarks that measure interpretability and trustworthiness directly to safeguard multimedia integrity in an age of pervasive AI-generated data.
Giulia Boato is full Professor at the Department of Information Engineerig and Computer Science (DISI) of the University of Trento (Italy). In 2006 she was visiting researcher at the Signal Theory and Communications Department of the University of Vigo (Spain). Since 2006 she has been collaborating with the Signal Processing Department of the Tampere University of Technology (Finland), in particular with prof. Karen Egiazarian. Since 2009 she has been working with prof. Hany Farid of the Dartmouth College (USA) on digital image and video forensics techniques.
Information Technologies Institute (ITI-CERTH), Greece
Course title: "Foundation models for video understanding tasks"
Abstract: In this talk we will focus on foundation models for video understanding, specifically Vision-Language Models (VLMs) and Video VLMs. We will discuss their potential applications, with special attention to Visual Question Answering (VQA), and the corresponding benchmarks. We will then identify key challenges associated with the use of such models, including computational and memory limitations arising from the temporal dimension of video streams. We will present in more detail selected methods for Video QA that are based on of-the-shelf VLMs but go beyond utilizing just the raw video frames in order to leverage additional context; and, scale in memory-efficient and latency-conscious ways. We will also present selected methods for improving the efficiency of the base VLMs using knowledge distillation.
Vasileios Mezaris is a Greek researcher and expert in multimedia understanding and artificial intelligence. He is currently a Research Director (Senior Researcher Grade A) at the Information Technologies Institute (ITI) of the Centre for Research and Technology Hellas (CERTH) in Thessaloniki, Greece, where he also leads the Intelligent Digital Transformation Laboratory. His research focuses on areas such as image and video analysis and annotation, machine learning and deep learning for multimedia understanding, big data analytics, multimedia indexing and retrieval, as well as explainable and green AI. He has served in editorial roles for major journals including IEEE Signal Processing Letters and the IEEE Transactions on Multimedia. Mezaris has participated in numerous research projects, often as coordinator or principal investigator. He is also a Senior Member of the IEEE.
University College London (UCL), UK
Course title: "Graph signal processing toward graph generative models"
Abstract: Learning to generate graphs is a compelling and fundamental challenge at the intersection of signal processing and machine learning. While graph generative models are increasingly used for synthesizing complex networks, examining them through a Graph Signal Processing (GSP) lens provides crucial structural guarantees and analytical insight. To establish this perspective, this talk first introduces the principles of generative models at large, transitions to generative modeling on graph-structured data, and ultimately examines graph generative models through a GSP lens—drawing upon a unified framework that bridges graph topology learning with graph generation.
Laura Toni is a Professor in the Department of Electronic and Electrical Engineering at University College London (UCL). She specializes in large-scale signal processing, machine learning, reinforcement learning, and multimedia systems. Toni is a Turing Fellow at the Alan Turing Institute and a member of European Laboratory for Learning and Intelligent Systems. She earned her Ph.D. in electrical engineering from the University of Bologna and held postdoctoral positions at University of California, San Diego and École Polytechnique Fédérale de Lausanne. Her work includes research on online adaptive strategies, graph processing, and dynamic network decision-making. She has (co-)authored numerous publications and serves in leadership roles for international conferences and journals.
Concordia University, Université de Montréal, Canada
Course title: "Compact audio and speech representations"
Abstract: Nowadays, compact representations, such as embeddings, discrete units, and low-dimensional acoustic features, are common to a variety of audio and speech-related applications. They enable efficient storage, faster transmission, and seamless integration with modern generative models. Historically, speech and audio processing relied on handcrafted parametric representations to efficiently compress and model data. The first part of this course aims to provide a quick overview of traditional methods for pre-processing speech and audio signals, as well as for their compression at low bit rates. The second part will cover modern discrete representations learned directly from raw speech and audio waveforms using neural networks. We will discuss the main approaches, their strengths and limitations, and their role in multimodal large language models.
Mirco Ravanelli is a Professor in the Department of Computer Science and Software Engineering at Concordia University. He works on deep learning for sequence processing and Conversational AI. He is also an Adjunct Professor at Université de Montréal and an Associate Member of the Mila – Quebec AI Institute, where he was a postdoctoral researcher under Yoshua Bengio . Ravanelli specializes in speech processing, machine learning, and representation learning. He is the creator and leader of SpeechBrain, an open-source toolkit for speech and conversational AI. His work has been recognized with awards such as the 2022 Amazon Research Award. His research contributes to advancing conversational systems and speech technologies.
FBK Trento, Italy, and Former Amazon Inc.
Course title part 2: "Compact audio and speech representations"
Abstract: Nowadays, compact representations, such as embeddings, discrete units, and low-dimensional acoustic features, are common to a variety of audio and speech-related applications. They enable efficient storage, faster transmission, and seamless integration with modern generative models. Historically, speech and audio processing relied on handcrafted parametric representations to efficiently compress and model data. The first part of this course aims to provide a quick overview of traditional methods for pre-processing speech and audio signals, as well as for their compression at low bit rates. The second part will cover modern discrete representations learned directly from raw speech and audio waveforms using neural networks. We will discuss the main approaches, their strengths and limitations, and their role in multimodal large language models.
Maurizio Omologo has worked in speech recognition, artificial intelligence, and microphone array signal processing. His research spans robust speech recognition, deep learning for far-field audio, and acoustic signal enhancement.
He has authored and co-authored many influential publications in automatic speech recognition and related fields.
He was historically affiliated with Fondazione Bruno Kessler (FBK) in Italy, where he worked on distant-speech interaction projects.
He has collaborated on projects involving transformer-based and neural speech models. He has also been involved in the development of practical corpora for speech data analysis. His work bridges speech technology and machine learning, influencing both theory and applications. Maurizio Omologo has worked at Amazon as a Principal Applied Scientist, particularly associated with Amazon Alexa / speech technology research.
École des Ponts ParisTech (ENPC), France
Course title: Generating 3D Human Motion with Language
Abstract: This lecture will describe works bridging natural language and 3D human motions. In particular, we will look at the evolution of text-driven generative models, e.g., given a text like `jump forward with your arms raised', can we synthesize a corresponding 3D human motion? This is a relatively recent field, which has witnessed a sudden growth. We will review some of the key works based on VAEs and diffusion models. There will be a special emphasis on compositionality to handle finegrained textual descriptions. In the last part, we will also look at results of follow-up works on text-to-motion retrieval (CLIP-like models for 3D motions), text-based motion editing, motion captioning, and 3D hand motion generation.
Gül Varol is a computer vision researcher and Associate Professor (permanent researcher) at École des Ponts ParisTech in France. Her work focuses on vision-and-language research, video understanding, human motion synthesis, and sign language analysis. She previously worked as a postdoctoral researcher at the University of Oxford (Visual Geometry Group) and received her PhD from Inria Paris and École Normale Supérieure (ENS).