Skip to content

Repository files navigation

English | 中文

📚 Video CNN/XAI Research Hub

Automated paper curation for video deep learning & explainability research

papers core strongly_related arXiv Semantic Scholar last_update license


Quick Navigation · 🏆 Influential · 🔥 Trending · 📄 Core · 📎 Strongly Related · 🏷️ Topics · 📈 Trends · 🖥️ Dashboard · 📋 Full List


📊 Overview

Metric Count
📚 Total Papers 894
🔥 Core Papers 379
📎 Strongly Related 515
🆕 New This Month 194
📡 arXiv 575
🔬 Semantic Scholar 314
🔗 CrossRef Enriched 10
✍️ Manual 5
⏰ Last Updated 2026-08-23 02:27:02

🏆 Top 5 Most Influential

This list highlights long-term impact; see Trending for recent work.

Rank Title Citations Score
1 Visualizing and Understanding Convolutional Networks 7502 5.4
2 Is Space-Time Attention All You Need for Video Understanding 3135 5.6
3 A Closer Look at Spatiotemporal Convolutions for Action Reco 3638 5.1
4 Convolutional Two-Stream Network Fusion for Video Action Rec 2771 5.0
5 A survey of methods for explaining Black Box Models 3548 4.1

🔥 Latest Trending (2024-2026)

Year Title Summary Citations Score
2024 VideoMamba: State Space Model for Efficient Video Understand Addressing the dual challenges of local redundancy and global dependencies in vi 551 4.7
2024 LongVU: Spatiotemporal Adaptive Compression for Long Video-L Multimodal Large Language Models (MLLMs) have shown promising progress in unders 314 4.7
2024 Benchmarking Micro-Action Recognition: Dataset, Methods, and Micro-action is an imperceptible non-verbal behaviour characterised by low-inten 138 4.0
2024 Isolated Video-Based Sign Language Recognition Using a Hybri Sign language is a complex language that uses hand gestures, body movements, and 54 4.7
2025 Video deepfake detection using a hybrid CNN-LSTM-Transformer The proliferation of deepfake technology poses significant challenges due to its 51 5.0

🔥 Latest Core Papers

Year Title Summary Author Score
2026 Addressable Memory for Video World Models We study visual persistence in interactive video world models. These models rely Xindi Wu, Sven Elflein+ 5.7
2026 Searching Videos as Trees: Self-Correcting Agents for Ground Grounded long-video question answering (Grounded LVQA) requires answering a ques Ce Zhang, Ziyang Wang+ 5.5
2026 GROVE: Growing and Reasoning over Temporally Stratified Memo A wearable assistant should both answer questions about its visual history and r Sitong Gong, Caixin Kang+ 5.5
2026 Video-DeepResearch: Towards the Next-Generation Multimodal D We introduce Video-DeepResearch (Video-DR), extending multimodal agents from sta Zhen Fang, Yu Zeng+ 5.5
2026 HelloWorld: Enabling Socially Interactive Characters in Vide Despite the remarkable recent progress of video world models, social interaction Liangyang Ouyang, Ruicong Liu+ 5.5
2026 Towards Expert-level Medical AI for Real-time Video Consulta Audio-visual interaction is the standard for patient-physician consultations, en Mahvish Nagda, Jihyeon Lee+ 5.5
2026 X-LMC: Cross-View Spatiotemporal Collateral Circulation Scor Digital subtraction angiography (DSA) is the reference standard for leptomeninge Maedeh Hafezi Moghadas, Hakim Baazaoui+ 5.5
2026 Visual Representation Matters: Exploiting Temporal Differenc Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by intro Zehua Chen, Junyou Wang+ 5.4
2026 Video Understanding: From Geometry and Semantics to Unified Video understanding aims to enable models to perceive, reason about, and interac Zhaochong An, Zirui Li+ 5.2
2026 Audio-Visual Flamingo: Open Audio-Visual Intelligence for Lo We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art au Sreyan Ghosh, Arushi Goel+ 5.2
2026 Sparse Evidence Can Suffice: Agentic Evidence Seeking for Mu Multimodal video misinformation detection is commonly formulated as a holistic v Haochen Zhao, Yongxiu Xu+ 5.2
2026 HAS: Highlight-guided Attention Steering for Multimodal LLM Video understanding has become more and more important with the growth of Artifi Rui Chu, Yingjie Lao 5.2
2026 Time-Reversed Imaging: A Multimodal Benchmark and Framework We introduce time-reversed imaging, a new paradigm that infers what just happene Jorge Bacca, Kebin Contreras+ 5.2
2026 Test-Time Adaptation via Dual Distillation for Videos Under Deep learning models have achieved state-of-the-art performance in several compu André Sacilotti, Samuel Felipe dos Santos+ 5.2
2026 CADER: Confidence-Aware Dynamic Evidence Reasoning for Long- Long-video understanding increasingly relies on large vision-language models and Jinlong Yang, Wenhao Zhang+ 5.2
2026 EgoPlay: Event-Triggered Video Editing for Egocentric Stream We introduce EgoPlay, an event-triggered video-to-video editor for egocentric st Jinjie Mai, Gordon Guocheng Qian+ 5.2
2026 Ripple: Real-Time Streaming Audio-Video Generation With Cros Audio-video generative models achieve impressive quality but suffer from high la Yanbo Ding, Zhizhi Guo+ 5.2
2026 EchoCache: Energy-Guided Cross-Modal Caching for Efficient A Audio-driven video generation (A2V) has achieved promising progress in synthesiz Jiayu Chen, Xiaoyu Wu+ 5.2
2026 Robust and Efficient Motion Reasoning for Privacy-Aware Clas Can computer vision help make classrooms safer? In this pilot study, we investig Paritosh Parmar, Landy Lan+ 5.2
2026 HOPE: Hand-Object Pressure Estimation from Monocular Videos Estimating physical pressure from vision is essential for understanding contact- Subin Jeon, Byungjun Kim+ 5.2

📎 Strongly Related Papers

Year Title Summary Author Score
2026 SM4RT: Learning Structured Motion Geometry for 4D Reconstruc Geometry Foundation Models (GFMs) have substantially advanced monocular 3D recon Shing Ho J. Lin, Wenzhao Zheng+ 3.9
2026 From Local Payoffs to Global Instabilities: A Spectral Carto We develop a motif-based framework for spatiotemporal chaos in spatial evolution Ozgur Aydogmus 3.9
2026 From Passive Video to Editable Experience: Physically Ground The key bottleneck in embodied AI is not model architecture but data. Although b Jia Luo 3.9
2026 QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Q Fault-tolerant quantum computing (FTQC) relies on quantum error correction to su Ran Miao, Rui Luo+ 3.9
2026 Faster-WAM: Do World Action Models Need Deep Action Modules? World Action Models (WAMs) couple robot action prediction with video world model Liheng Ma, Rui Heng Yang+ 3.9
2026 Context-Aware Mixture of Domain Experts for Bodily Expressio The same body posture can convey entirely different emotions depending on its su Mohammad Mahdi Dehshibi, David Masip 3.9
2026 Identity-Faithful Audio-Visual Target Speaker Extraction wit Audio-visual target speaker extraction should return the speaker indicated by th Peijun Yang, Zhan Jin+ 3.9
2026 Multimodal Spatiotemporal Atmospheric Data Assimilation with Data assimilation (DA) uses Bayesian inference to update the state of a numerica Dibyajyoti Chakraborty, Romit Maulik 3.9
2026 BendTwin: Robust Dense-to-Sparse Physical Reconstruction wit Reconstructing objects with mechanical properties from video observations enable Yixiong Jing, Qi Wang+ 3.9
2026 Geometry-Aware Camera Localization for Bronchoscopy Camera localization in bronchoscopy remains a challenging problem due to stringe Lumin Chen, Qingyao Tian+ 3.9

📅 2026 (399 papers)
Tag Title Summary Author Score
🔥 Addressable Memory for Video World Models We study visual persistence in interactive video world models. These m Xindi Wu, Sven Elflein+ 5.7
🔥 Searching Videos as Trees: Self-Correcting Agents Grounded long-video question answering (Grounded LVQA) requires answer Ce Zhang, Ziyang Wang+ 5.5
🔥 GROVE: Growing and Reasoning over Temporally Strat A wearable assistant should both answer questions about its visual his Sitong Gong, Caixin Kang+ 5.5
🔥 Video-DeepResearch: Towards the Next-Generation Mu We introduce Video-DeepResearch (Video-DR), extending multimodal agent Zhen Fang, Yu Zeng+ 5.5
🔥 HelloWorld: Enabling Socially Interactive Characte Despite the remarkable recent progress of video world models, social i Liangyang Ouyang, Ruicong Liu+ 5.5
🔥 Towards Expert-level Medical AI for Real-time Vide Audio-visual interaction is the standard for patient-physician consult Mahvish Nagda, Jihyeon Lee+ 5.5
🔥 X-LMC: Cross-View Spatiotemporal Collateral Circul Digital subtraction angiography (DSA) is the reference standard for le Maedeh Hafezi Moghadas, Hakim Baazaoui+ 5.5
🔥 Visual Representation Matters: Exploiting Temporal Video-to-audio (V2A) generation extends image-to-audio generation (I2A Zehua Chen, Junyou Wang+ 5.4
🔥 Video Understanding: From Geometry and Semantics t Video understanding aims to enable models to perceive, reason about, a Zhaochong An, Zirui Li+ 5.2
🔥 Audio-Visual Flamingo: Open Audio-Visual Intellige We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of- Sreyan Ghosh, Arushi Goel+ 5.2
🔥 Sparse Evidence Can Suffice: Agentic Evidence Seek Multimodal video misinformation detection is commonly formulated as a Haochen Zhao, Yongxiu Xu+ 5.2
🔥 HAS: Highlight-guided Attention Steering for Multi Video understanding has become more and more important with the growth Rui Chu, Yingjie Lao 5.2

Showing 12 of 399 papers. See ALL_PAPERS.md for all entries.

📅 2025 (106 papers)
Tag Title Summary Author Score
🔥 Enhancing Video Understanding: Deep Neural Network It's no secret that video has become the primary way we share informat Amir Hosein Fadaei, Mohammad-Reza A. Dehaqani 5.8
🔥 Fine tuning 3D Convolutional Networks for enhanced The study of Human Activity Recognition (HAR) has attracted considerab Abir Frad, Hend Basly+ 5.4
🔥 A Hybrid 3D CNNs Transformer Architecture for Vide Video-Based Human Action Recognition (HAR) remains challenging due to Engin Seven, Eylem Yücel Demirel 5.4
🔥 Video-CoT: A Comprehensive Dataset for Spatiotempo Video content comprehension is essential for various applications, ran Shuyi Zhang, Xiaoshuai Hao+ 5.2
🔥 V-STaR: Benchmarking Video-LLMs on Video Spatio-Te Human processes video reasoning in a sequential spatio-temporal reason Zixu Cheng, Jian Hu+ 5.2
🔥 Video deepfake detection using a hybrid CNN-LSTM-T The proliferation of deepfake technology poses significant challenges G. Petmezas, Vazgken Vanian+ 5.0
🔥 TinyLLaVA-Video: Towards Smaller LMMs for Video Un Video behavior recognition and scene understanding are fundamental tas Xingjian Zhang, Xi Weng+ 4.9
🔥 Harnessing Synthetic Preference Data for Enhancing While Video Large Language Models (Video-LLMs) have demonstrated remar Sameep Vani, Shreyas Jena+ 4.9
🔥 AceVFI: A Comprehensive Survey of Advances in Vide Video Frame Interpolation (VFI) is a core low-level vision task that s Dahyeon Kye, Changhyun Roh+ 4.9
🔥 How Much 3D Do Video Foundation Models Encode? Videos are continuous 2D projections of 3D worlds. After training on l Zixuan Huang, Xiang Li+ 4.9
🔥 A Novel 3D Convolutional Neural Network-Based Deep Accurate analysis of medical videos remains a major challenge in deep M. K. Dhar, Mou Deb+ 4.9
🔥 RepAttn3D: Re-parameterizing 3D attention with spa The technique of structural re-parameterization has been widely adopte Xiusheng Lu, Lechao Cheng+ 4.8

Showing 12 of 106 papers. See ALL_PAPERS.md for all entries.

📅 2024 (103 papers)
Tag Title Summary Author Score
🔥 InternVideo2: Scaling Foundation Models for Multim We introduce InternVideo2, a new family of video foundation models (Vi Yi Wang, Kunchang Li+ 5.2
🔥 Various frameworks for integrating image and video Human action recognition has been identified as an important research Shaimaa Yosry, Lamiaa A. Elrefaei+ 4.8
🔥 Video-based Exercise Classification and Activated This paper introduces a simple yet effective strategy for exercise cla Manvik Pasula, Pramit Saha 4.7
🔥 LongVU: Spatiotemporal Adaptive Compression for Lo Multimodal Large Language Models (MLLMs) have shown promising progress Xiaoqian Shen, Yunyang Xiong+ 4.7
🔥 VideoMamba: State Space Model for Efficient Video Addressing the dual challenges of local redundancy and global dependen Kunchang Li, Xinhao Li+ 4.7
🔥 Isolated Video-Based Sign Language Recognition Usi Sign language is a complex language that uses hand gestures, body move Diksha Kumari, Radhey Shyam Anand 4.7
🔥 Can VLMs be used on videos for action recognition? Recent advancements have introduced multiple vision-language models (V Harsh Lunia 4.4
🔥 Prompting Video-Language Foundation Models with Do Video Question Answering (VideoQA) represents a crucial intersection b Ting Yu, Kunhao Fu+ 4.4
🔥 Relevance-guided Audio Visual Fusion for Video Sal Audio data, often synchronized with video frames, plays a crucial role Li Yu, Xuanzhe Sun+ 4.4
🔥 Automated diagnosis of respiratory diseases from l An automated computerized approach can aid radiologists in the early d Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan+ 4.3
🔥 Facial Expression Recognition in Video Using 3D-CN The focus of research work presented in this paper on improving perfor Sathisha G, C. K. Subbaraya+ 4.3
🔥 Interpretability in Video-based Human Action Recog Interpretability plays a vital role in understanding complex deep lear Jorge Garcia-Torres Fernandez 4.3

Showing 12 of 103 papers. See ALL_PAPERS.md for all entries.

📅 2023 (57 papers)
Tag Title Summary Author Score
🔥 Hierarchical Spatiotemporal Feature Fusion Network Current video saliency prediction methods have made great progress rel Yunzuo Zhang, Tian Zhang+ 5.7
🔥 Video-FocalNets: Spatio-Temporal Focal Modulation Recent video recognition models utilize Transformer models for long-ra Syed Talal Wasim, Muhammad Uzair Khattak+ 5.5
🔥 Deep Neural Networks in Video Human Action Recogni Currently, video behavior recognition is one of the most foundational Zihan Wang, Yang Yang+ 5.5
🔥 Video Understanding with Large Language Models: A With the burgeoning growth of online video platforms and the escalatin Yolo Y. Tang, Jing Bi+ 5.2
🔥 Audio-visual Saliency for Omnidirectional Videos Visual saliency prediction for omnidirectional videos (ODVs) has shown Yuxin Zhu, Xilei Zhu+ 5.2
🔥 Understanding Video Transformers for Segmentation: Video segmentation encompasses a wide range of categories of problem f Rezaul Karim, Richard P. Wildes 4.8
🔥 VMC: Video Motion Customization using Temporal Att Text-to-video diffusion models have advanced video generation signific Hyeonho Jeong, Geon Yeong Park+ 4.7
🔥 A Video Is Worth 4096 Tokens: Verbalize Videos To Multimedia content, such as advertisements and story videos, exhibit a Aanisha Bhattacharya, Yaman K Singla+ 4.6
🔥 A dynamic gesture recognition method based on R(2+ Efficient spatial-temporal feature extraction from input video streams Yupeng Huo, Jie Shen+ 4.6
🔥 Spatio-Temporal Features based Human Action Recogn —Recognition of human intention is crucial and challenging due to subt Saifuddin Saif, E. Wollega+ 4.5
🔥 AMS-Net: Modeling Adaptive Multi-Granularity Spati Effective spatio-temporal modeling as a core of video representation l Qilong Wang, Qiyao Hu+ 4.5
🔥 Video Traffic Analysis for Real-Time Emotion Recog Since the outbreak of the COVID-19 crisis, the transition to remote ed Ayoub Sassi, W. Jaafar+ 4.5

Showing 12 of 57 papers. See ALL_PAPERS.md for all entries.

📅 2022 (51 papers)
Tag Title Summary Author Score
🔥 Large-scale Robustness Analysis of Video Action Re We have seen a great progress in video action recognition in recent ye Madeline Chantry Schiappa, Naman Biyani+ 5.5
🔥 VRT: A Video Restoration Transformer Video restoration (e. g. , video super-resolution) aims to restore hig Jingyun Liang, Jiezhang Cao+ 5.5
🔥 3D Convolutional with Attention for Action Recogni Human action recognition is one of the challenging tasks in computer v Labina Shrestha, Shikha Dubey+ 5.0
🔥 UniFormerV2: Spatiotemporal Learning by Arming Ima Learning discriminative spatiotemporal representation is the key probl Kunchang Li, Yali Wang+ 5.0
🔥 Video Human Action Recognition Algorithm Based on The traditional action recognition algorithm based on manual feature e Yu Wang, Jiaxi Sun 4.8
🔥 No-Reference Video Quality Assessment Using Multi- With the constantly growing popularity of video-based services and app D. Varga 4.8
🔥 Action Recognition Using Action Sequences Optimiza Effective extraction and representation of action information are crit Xin Xiong, Weidong Min+ 4.5
🔥 Enhancing Deformable Convolution based Video Frame This paper presents a new deformable convolution-based video frame int Duolikun Danier, Fan Zhang+ 4.4
🔥 Skeleton Graph-Neural-Network-Based Human Action R Human action recognition has been applied in many fields, such as vide Miao Feng, Jean Meunier 4.4
🔥 Two-stream fusion model using 3D-CNN and 2D-CNN vi Hand gestures are useful tools for many applications in the human-comp Debajit Sarma, V. Kavyasree+ 4.3
🔥 Sign Language Recognition Based on R(2+1)D With Sp Previous work utilized three-dimensional (3-D) convolutional neural ne Xiangzu Han, Fei Lu+ 4.3
🔥 Video Visual Relation Detection via 3D Convolution Video visual relation detection, which aims to detect the visual relat Mingcheng Qu, Jianxun Cui+ 4.3

Showing 12 of 51 papers. See ALL_PAPERS.md for all entries.

📅 2021 (49 papers)
Tag Title Summary Author Score
🔥 Spatiotemporal Dilated Convolution with Uncertain In this paper, we propose a novel SpatioTemporal convolutional Dense N Yu-Jen Ma, Hong-Han Shuai+ 6.3
🔥 Action Transformer: A Self-Attention Model for Sho Deep neural networks based purely on attention have been successful ac Vittorio Mazzia, Simone Angarano+ 5.7
🔥 Temporal-attentive Covariance Pooling Networks for For video recognition task, a global representation summarizing the wh Zilin Gao, Qilong Wang+ 5.7
🔥 Is Space-Time Attention All You Need for Video Und We present a convolution-free approach to video classification built e Gedas Bertasius, Heng Wang+ 5.6
🔥 TAda! Temporally-Adaptive Convolutions for Video U Spatial convolutions are widely used in numerous deep video models. It Ziyuan Huang, Shiwei Zhang+ 5.2
🔥 Efficient Action Recognition with Introducing R(2+ The mainstream methods in video action recognition includes 3D convolu Hao Jin, Jianming Yang+ 5.2
🔥 Token Shift Transformer for Video Classification Transformer achieves remarkable successes in understanding 1 and 2-dim Hao Zhang, Y. Hao+ 5.0
🔥 Improved CNN-based Learning of Interpolation Filte The versatility of recent machine learning approaches makes them ideal Luka Murn, Saverio Blasi+ 4.9
🔥 SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-Ki This report presents the technical details of our submission to the EP Swathikiran Sudhakaran, Adrian Bulat+ 4.7
🔥 "Knights": First Place Submission for VIPriors21 A This technical report presents our approach "Knights" to solve the act Ishan Dave, Naman Biyani+ 4.4
🔥 Revisiting Video Saliency Prediction in the Deep L Predicting where people look in static scenes, a. k. a visual saliency Wenguan Wang, Jianbing Shen+ 4.3
🔥 Recent Advances in Video Action Recognition with 3 SUMMARY The performance of video action recognition has improved signi Kensho Hara 4.3

Showing 12 of 49 papers. See ALL_PAPERS.md for all entries.

📅 2020 (41 papers)
Tag Title Summary Author Score
🔥 Deep Analysis of CNN-based Spatio-temporal Represe In recent years, a number of approaches based on 2D or 3D convolutiona Chun-Fu Chen, Rameswar Panda+ 5.8
🔥 Unified Image and Video Saliency Modeling Visual saliency modeling for images and videos is treated as two indep Richard Droste, Jianbo Jiao+ 5.7
🔥 TAM: Temporal Adaptive Module for Video Recognitio Video data is with complex temporal dynamics due to various factors su Zhaoyang Liu, Limin Wang+ 5.7
🔥 Dissected 3D CNNs: Temporal Skip Connections for E Convolutional Neural Networks with 3D kernels (3D-CNNs) currently achi Okan Köpüklü, Stefan Hörmann+ 5.5
🔥 TEA: Temporal Excitation and Aggregation for Actio Temporal modeling is key for action recognition in videos. It normally Yan Li, Bin Ji+ 5.5
🔥 RANP: Resource Aware Neuron Pruning at Initializat Although 3D Convolutional Neural Networks (CNNs) are essential for mos Zhiwei Xu, Thalaiyasingam Ajanthan+ 5.2
🔥 Developing Motion Code Embedding for Action Recogn In this work, we propose a motion embedding strategy known as motion c Maxat Alibayev, David Paulius+ 5.2
🔥 Would Mega-scale Datasets Further Enhance Spatiote How can we collect and use a video dataset to further improve spatiote Hirokatsu Kataoka, Tenga Wakamiya+ 5.0
🔥 Learnable Sampling 3D Convolution for Video Enhanc A key challenge in video enhancement and action recognition is to fuse Shuyang Gu, Jianmin Bao+ 5.0
🔥 Challenge report:VIPriors Action Recognition Chall This paper is a brief report to our submission to the VIPriors Action Zhipeng Luo, Dawei Xu+ 4.7
🔥 Res3ATN -- Deep 3D Residual Attention Network for Hand gesture recognition is a strenuous task to solve in videos. In th Naina Dhingra, Andreas Kunz 4.6
🔥 Toward Accurate Person-level Action Recognition in Detecting and recognizing human action in videos with crowded scenes i Li Yuan, Yichen Zhou+ 4.4

Showing 12 of 41 papers. See ALL_PAPERS.md for all entries.

📅 2019 (27 papers)
Tag Title Summary Author Score
🔥 Spatio-Temporal FAST 3D Convolutions for Human Act Effective processing of video input is essential for the recognition o Alexandros Stergiou, Ronald Poppe 5.8
🔥 A review of Convolutional-Neural-Network-based act Abstract Video action recognition is widely applied in video indexing, Guangle Yao, Tao Lei+ 5.3
🔥 Spatiotemporal distilled dense-connectivity networ Abstract Two-stream convolutional neural networks show great promise f Wangli Hao, Zhaoxiang Zhang 5.3
🔥 Resource Efficient 3D Convolutional Neural Network Recently, convolutional neural networks with 3D kernels (3D CNNs) have Okan Köpüklü, Neslihan Kose+ 5.0
🔥 Image and Video Compression with Neural Networks: In recent years, the image and video coding technologies have advanced Siwei Ma, Xinfeng Zhang+ 4.9
🔥 Improving Action Recognition with the Graph-Neural Recent human action recognition methods mainly model a two-stream or 3 Wu Luo, Chongyang Zhang+ 4.5
🔥 Multi-teacher Knowledge Distillation for Compresse Recently, convolutional neural networks (CNNs) have seen great progres Meng-Chieh Wu, C. Chiu+ 4.2
🔥 Motion Sickness Prediction in Stereoscopic Videos In this paper, we propose a three-dimensional (3D) convolutional neura Tae Min Lee, Jong-Chul Yoon+ 4.2
🔥 Predicting 3D Human Dynamics from Video Given a video of a person in action, we can easily guess the 3D future Jason Y. Zhang, Panna Felsen+ 4.1
🔥 FBK-HUPBA Submission to the EPIC-Kitchens 2019 Act In this report we describe the technical details of our submission to Swathikiran Sudhakaran, Sergio Escalera+ 4.1
📎 Explainable Deep Learning for Video Recognition Ta The popularity of Deep Learning for real-world applications is ever-gr Liam Hiley, Alun Preece+ 3.9
📎 Deep 3D Convolutional Neural Network for Automated Computer Aided Diagnosis has emerged as an indispensible technique for Sumita Mishra, Naresh Kumar Chaudhary+ 3.9

Showing 12 of 27 papers. See ALL_PAPERS.md for all entries.

📅 2018 (28 papers)
Tag Title Summary Author Score
🔥 Interpretable Spatio-temporal Attention for Video Inspired by the observation that humans are able to process videos eff Lili Meng, Bo Zhao+ 7.4
🔥 Revisiting Video Saliency: A Large-scale Benchmark In this work, we contribute to video saliency research in two ways. Fi Wenguan Wang, Jianbing Shen+ 5.3
🔥 Review of Visual Saliency Detection with Comprehen Visual saliency detection model simulates the human visual system to p Runmin Cong, Jianjun Lei+ 5.0
🔥 Reduced-Gate Convolutional LSTM Using Predictive C Spatiotemporal sequence prediction is an important problem in deep lea Nelly Elsayed, Anthony S. Maida+ 4.9
🔥 Recurrent Convolutions for Causal 3D CNNs Recently, three dimensional (3D) convolutional neural networks (CNNs) Gurkirt Singh, Fabio Cuzzolin 4.8
🔥 Non-local NetVLAD Encoding for Video Classificatio This paper describes our solution for the 2$^\text{nd}$ YouTube-8M vid Yongyi Tang, Xing Zhang+ 4.6
🔥 Morph: Flexible Acceleration for 3D CNN-Based Vide The past several years have seen both an explosion in the use of Convo Kartik Hegde, R. Agrawal+ 4.5
🔥 ECO: Efficient Convolutional Network for Online Vi The state of the art in video understanding suffers from two problems: Mohammadreza Zolfaghari, Kamaljeet Singh+ 4.4
🔥 Non-Local Video Denoising by CNN Non-local patch based methods were until recently state-of-the-art for Axel Davy, Thibaud Ehret+ 4.4
🔥 Recurrence to the Rescue: Towards Causal Spatiotem Recently, three dimensional (3D) convolutional neural networks (CNNs) Gurkirt Singh, Fabio Cuzzolin 4.3
🔥 Benchmark 3D eye-tracking dataset for visual salie Visual Attention Models (VAMs) predict the location of an image or vid Amin Banitalebi-Dehkordi, Eleni Nasiopoulos+ 4.2
🔥 SlowFast Networks for Video Recognition We present SlowFast networks for video recognition. Our model involves Christoph Feichtenhofer, Haoqi Fan+ 4.2

Showing 12 of 28 papers. See ALL_PAPERS.md for all entries.

📅 2017 (19 papers)
Tag Title Summary Author Score
🔥 The Monkeytyping Solution to the YouTube-8M Video This article describes the final solution of team monkeytyping, who fi He-Da Wang, Teng Zhang+ 5.2
🔥 A Closer Look at Spatiotemporal Convolutions for A In this paper we discuss several forms of spatiotemporal convolutions Du Tran, Heng Wang+ 5.1
🔥 Predicting Video Saliency with Object-to-Motion CN Over the past few years, deep neural networks (DNNs) have exhibited gr Lai Jiang, Mai Xu+ 5.0
🔥 Hierarchical Deep Recurrent Architecture for Video This paper introduces the system we developed for the Youtube-8M Video Luming Tang, Boyang Deng+ 4.9
🔥 Graph-Theoretic Spatiotemporal Context Modeling fo As an important and challenging problem in computer vision, video sali Lina Wei, Fangfang Wang+ 4.8
🔥 Video Classification With CNNs: Using The Codec As We investigate video classification via a two-stream convolutional neu Aaron Chadha, Alhabib Abbas+ 4.7
🔥 Facial Expression Recognition Using Enhanced Deep Deep Neural Networks (DNNs) have shown to outperform traditional metho Behzad Hasani, Mohammad H. Mahoor 4.4
🔥 Temporal Relational Reasoning in Videos Temporal relational reasoning, the ability to link meaningful transfor Bolei Zhou, A. Andonian+ 4.2
🔥 Two-Stream 3D Convolutional Neural Network for Ske It remains a challenge to efficiently extract spatialtemporal informat Hong Liu, Juanhui Tu+ 4.2
🔥 Spatio-Temporal Facial Expression Recognition Usin Automated Facial Expression Recognition (FER) has been a challenging t Behzad Hasani, Mohammad H. Mahoor 4.1
📎 A Brief Survey of Deep Reinforcement Learning Deep reinforcement learning is poised to revolutionise the field of AI Kai Arulkumaran, Marc Peter Deisenroth+ 3.9
📎 Rethinking Spatiotemporal Feature Learning For Vid Rethinking Spatiotemporal Feature Learning For Video Understanding Saining Xie, Chen Sun+ 3.9

Showing 12 of 19 papers. See ALL_PAPERS.md for all entries.

📅 2016 (6 papers)
Tag Title Summary Author Score
🔥 Convolutional Two-Stream Network Fusion for Video Recent applications of Convolutional Neural Networks (ConvNets) for hu Christoph Feichtenhofer, A. Pinz+ 5.0
🔥 Large-Scale Shape Retrieval with Sparse 3D Convolu In this paper we present results of performance evaluation of S3DCNN - Alexandr Notchenko, Ermek Kapushev+ 4.4
📎 Deep Learning for Saliency Prediction in Natural V The purpose of this paper is the detection of salient areas in natural Souad Chaabouni, Jenny Benois-Pineau+ 3.4
📎 Grad-CAM: Visual Explanations from Deep Networks v We propose a technique for producing ‘visual explanations’ for decisio Ramprasaath R. Selvaraju, Abhishek Das+ 3.3
📎 A Deep Learning Approach for Joint Video Frame and Reinforcement learning is concerned with identifying reward-maximizing Felix Leibfried, Nate Kushman+ 2.6
📎 Transfer learning with deep networks for saliency Transfer learning with deep networks for saliency prediction in natura S. Chaabouni, J. Benois-Pineau+ 2.6
📅 2015 (6 papers)
Tag Title Summary Author Score
🔥 C3D: Generic Features for Video Analysis We propose a simple yet effective approach for spatiotemporal feature Du Tran, Lubomir Bourdev+ 5.7
🔥 Intra-and-Inter-Constraint-based Video Enhancement Video enhancement plays an important role in various video application Yuanzhe Chen, Weiyao Lin+ 4.1
📎 Activity Recognition Using A Combination of Catego This paper presents a novel approach for automatic recognition of huma Weiyao Lin, Ming-Ting Sun+ 3.8
📎 A new network-based algorithm for human activity r In this paper, a new network-transmission-based (NTB) algorithm is pro Weiyao Lin, Yuanzhe Chen+ 3.8
📎 Group Event Detection with a Varying Number of Gro This paper presents a novel approach for automatic recognition of grou Weiyao Lin, Ming-Ting Sun+ 2.9
📎 VoxNet: A 3D Convolutional Neural Network for real VoxNet: A 3D Convolutional Neural Network for real-time object recogni D. Maturana, S. Scherer 2.8
📅 2014 (2 papers)
Tag Title Summary Author Score
🔥 Visualizing and Understanding Convolutional Networ Large Convolutional Network models have recently demonstrated impressi Matthew D. Zeiler, Rob Fergus 5.4
📎 Deep Inside Convolutional Networks: Visualising Im This paper addresses the visualisation of image classification models, Karen Simonyan, Andrea Vedaldi+ 3.7

📄 Full paper list: ALL_PAPERS.md

🏗️ Architecture

┌─────────────────────────────────────────────────────────────┐
│                    search_config.json                       │
│              (24 queries × 3 depth layers)                  │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                     Collector                               │
│         arXiv API  +  Semantic Scholar API                  │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│              Normalizer + CrossRef Enrichment               │
│        ID/version normalization, dedup, citations           │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                      Scorer                                 │
│    Keyword Match + Citations + Venue + Survey Bonus          │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                      Storage                                │
│   papers/index.jsonl + papers/quarantine.jsonl              │
└───────────────────────────┬─────────────────────────────────┘
                            │
          ┌─────────────────┼─────────────────┐
          ▼                 ▼                 ▼
    ┌──────────┐     ┌──────────┐     ┌──────────┐
    │ README   │     │ Feishu   │     │ Dashboard │
    │ (Display)│     │ (Notify) │     │  (HTML)   │
    └──────────┘     └──────────┘     └──────────┘

✨ Features

🎯 Smart Search 📊 Data Enhancement 🌐 Multi-Source 🔔 Auto Notify
Daily arXiv search CrossRef enrichment for new papers arXiv + Semantic Scholar Feishu Webhook
24 layered queries CN/EN summary generation Title dedup + ID norm Success/Failure alerts
Score-based filtering Topic clustering Citation + Venue boost GitHub Actions

⚙️ Auto Update

GitHub Actions aggregates arXiv and Semantic Scholar daily, then normalizes, deduplicates, scores, and enriches accepted papers with CrossRef.

📄 License

For academic research use only

About

视频 CNN 可解释性相关论文自动搜索与整理

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages