English | 中文
Automated paper curation for video deep learning & explainability research
Quick Navigation · 🏆 Influential · 🔥 Trending · 📄 Core · 📎 Strongly Related · 🏷️ Topics · 📈 Trends · 🖥️ Dashboard · 📋 Full List
| Metric | Count |
|---|---|
| 📚 Total Papers | 747 |
| 🔥 Core Papers | 302 |
| 📎 Strongly Related | 445 |
| 🆕 New This Month | 50 |
| 📡 arXiv | 431 |
| 🔬 Semantic Scholar | 311 |
| 🔗 CrossRef Enriched | 6 |
| ✍️ Manual | 5 |
| ⏰ Last Updated | 2026-08-07 03:55:43 |
This list highlights long-term impact; see Trending for recent work.
| Rank | Title | Citations | Score |
|---|---|---|---|
| 1 | Visualizing and Understanding Convolutional Networks | 7502 | 5.4 |
| 2 | Is Space-Time Attention All You Need for Video Understanding | 3135 | 5.6 |
| 3 | A Closer Look at Spatiotemporal Convolutions for Action Reco | 3638 | 5.1 |
| 4 | Convolutional Two-Stream Network Fusion for Video Action Rec | 2771 | 5.0 |
| 5 | A survey of methods for explaining Black Box Models | 3548 | 4.1 |
| Year | Title | Summary | Citations | Score |
|---|---|---|---|---|
| 2024 | VideoMamba: State Space Model for Efficient Video Understand | Addressing the dual challenges of local redundancy and global dependencies in vi | 551 | 4.7 |
| 2024 | LongVU: Spatiotemporal Adaptive Compression for Long Video-L | Multimodal Large Language Models (MLLMs) have shown promising progress in unders | 314 | 4.7 |
| 2024 | Benchmarking Micro-Action Recognition: Dataset, Methods, and | Micro-action is an imperceptible non-verbal behaviour characterised by low-inten | 138 | 4.0 |
| 2024 | Isolated Video-Based Sign Language Recognition Using a Hybri | Sign language is a complex language that uses hand gestures, body movements, and | 54 | 4.7 |
| 2025 | Video deepfake detection using a hybrid CNN-LSTM-Transformer | The proliferation of deepfake technology poses significant challenges due to its | 51 | 5.0 |
| Year | Title | Summary | Author | Score |
|---|---|---|---|---|
| 2026 | Searching Videos as Trees: Self-Correcting Agents for Ground | Grounded long-video question answering (Grounded LVQA) requires answering a ques | Ce Zhang, Ziyang Wang+ | 5.5 |
| 2026 | GROVE: Growing and Reasoning over Temporally Stratified Memo | A wearable assistant should both answer questions about its visual history and r | Sitong Gong, Caixin Kang+ | 5.5 |
| 2026 | Video-DeepResearch: Towards the Next-Generation Multimodal D | We introduce Video-DeepResearch (Video-DR), extending multimodal agents from sta | Zhen Fang, Yu Zeng+ | 5.5 |
| 2026 | HelloWorld: Enabling Socially Interactive Characters in Vide | Despite the remarkable recent progress of video world models, social interaction | Liangyang Ouyang, Ruicong Liu+ | 5.5 |
| 2026 | Visual Representation Matters: Exploiting Temporal Differenc | Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by intro | Zehua Chen, Junyou Wang+ | 5.4 |
| 2026 | Video Understanding: From Geometry and Semantics to Unified | Video understanding aims to enable models to perceive, reason about, and interac | Zhaochong An, Zirui Li+ | 5.2 |
| 2026 | Audio-Visual Flamingo: Open Audio-Visual Intelligence for Lo | We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art au | Sreyan Ghosh, Arushi Goel+ | 5.2 |
| 2026 | Sparse Evidence Can Suffice: Agentic Evidence Seeking for Mu | Multimodal video misinformation detection is commonly formulated as a holistic v | Haochen Zhao, Yongxiu Xu+ | 5.2 |
| 2026 | HAS: Highlight-guided Attention Steering for Multimodal LLM | Video understanding has become more and more important with the growth of Artifi | Rui Chu, Yingjie Lao | 5.2 |
| 2026 | Time-Reversed Imaging: A Multimodal Benchmark and Framework | We introduce time-reversed imaging, a new paradigm that infers what just happene | Jorge Bacca, Kebin Contreras+ | 5.2 |
| 2026 | Test-Time Adaptation via Dual Distillation for Videos Under | Deep learning models have achieved state-of-the-art performance in several compu | André Sacilotti, Samuel Felipe dos Santos+ | 5.2 |
| 2026 | CADER: Confidence-Aware Dynamic Evidence Reasoning for Long- | Long-video understanding increasingly relies on large vision-language models and | Jinlong Yang, Wenhao Zhang+ | 5.2 |
| 2026 | EgoPlay: Event-Triggered Video Editing for Egocentric Stream | We introduce EgoPlay, an event-triggered video-to-video editor for egocentric st | Jinjie Mai, Gordon Guocheng Qian+ | 5.2 |
| 2026 | Ripple: Real-Time Streaming Audio-Video Generation With Cros | Audio-video generative models achieve impressive quality but suffer from high la | Yanbo Ding, Zhizhi Guo+ | 5.2 |
| 2026 | EchoCache: Energy-Guided Cross-Modal Caching for Efficient A | Audio-driven video generation (A2V) has achieved promising progress in synthesiz | Jiayu Chen, Xiaoyu Wu+ | 5.2 |
| 2026 | Robust and Efficient Motion Reasoning for Privacy-Aware Clas | Can computer vision help make classrooms safer? In this pilot study, we investig | Paritosh Parmar, Landy Lan+ | 5.2 |
| 2026 | HOPE: Hand-Object Pressure Estimation from Monocular Videos | Estimating physical pressure from vision is essential for understanding contact- | Subin Jeon, Byungjun Kim+ | 5.2 |
| 2026 | IVEX-WA and IVEX-MetaStack Ensemble Models: A Transfer Learn | Human action recognition (HAR) using deep learning approaches has significantly | Md Tasnim Alam, Subhram Dasgupta+ | 5.0 |
| 2026 | Hierarchical Denoising For Multi-Step Visual Reasoning | Video models are evolving into vision foundation models, yet they still lack hum | Zezhong Qian, Xiaowei Chi+ | 5.0 |
| 2026 | Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision | Perceiving multimodal cues and forecasting fine-grained actions from an egocentr | Zhaofeng Shi, Heqian Qiu+ | 5.0 |
| Year | Title | Summary | Author | Score |
|---|---|---|---|---|
| 2026 | SM4RT: Learning Structured Motion Geometry for 4D Reconstruc | Geometry Foundation Models (GFMs) have substantially advanced monocular 3D recon | Shing Ho J. Lin, Wenzhao Zheng+ | 3.9 |
| 2026 | From Local Payoffs to Global Instabilities: A Spectral Carto | We develop a motif-based framework for spatiotemporal chaos in spatial evolution | Ozgur Aydogmus | 3.9 |
| 2026 | From Passive Video to Editable Experience: Physically Ground | The key bottleneck in embodied AI is not model architecture but data. Although b | Jia Luo | 3.9 |
| 2026 | QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Q | Fault-tolerant quantum computing (FTQC) relies on quantum error correction to su | Ran Miao, Rui Luo+ | 3.9 |
| 2026 | Faster-WAM: Do World Action Models Need Deep Action Modules? | World Action Models (WAMs) couple robot action prediction with video world model | Liheng Ma, Rui Heng Yang+ | 3.9 |
| 2026 | Context-Aware Mixture of Domain Experts for Bodily Expressio | The same body posture can convey entirely different emotions depending on its su | Mohammad Mahdi Dehshibi, David Masip | 3.9 |
| 2026 | Identity-Faithful Audio-Visual Target Speaker Extraction wit | Audio-visual target speaker extraction should return the speaker indicated by th | Peijun Yang, Zhan Jin+ | 3.9 |
| 2026 | Multimodal Spatiotemporal Atmospheric Data Assimilation with | Data assimilation (DA) uses Bayesian inference to update the state of a numerica | Dibyajyoti Chakraborty, Romit Maulik | 3.9 |
| 2026 | BendTwin: Robust Dense-to-Sparse Physical Reconstruction wit | Reconstructing objects with mechanical properties from video observations enable | Yixiong Jing, Qi Wang+ | 3.9 |
| 2026 | PSAM: Parameter-Free Spatiotemporal Attention Mechanism for | Spatiotemporal attention learning has always been a challenging research task in | Fuwei Zhang, Ruomei Wang+ | 3.8 |
📅 2026 (253 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Searching Videos as Trees: Self-Correcting Agents | Grounded long-video question answering (Grounded LVQA) requires answer | Ce Zhang, Ziyang Wang+ | 5.5 |
| 🔥 | GROVE: Growing and Reasoning over Temporally Strat | A wearable assistant should both answer questions about its visual his | Sitong Gong, Caixin Kang+ | 5.5 |
| 🔥 | Video-DeepResearch: Towards the Next-Generation Mu | We introduce Video-DeepResearch (Video-DR), extending multimodal agent | Zhen Fang, Yu Zeng+ | 5.5 |
| 🔥 | HelloWorld: Enabling Socially Interactive Characte | Despite the remarkable recent progress of video world models, social i | Liangyang Ouyang, Ruicong Liu+ | 5.5 |
| 🔥 | Visual Representation Matters: Exploiting Temporal | Video-to-audio (V2A) generation extends image-to-audio generation (I2A | Zehua Chen, Junyou Wang+ | 5.4 |
| 🔥 | Video Understanding: From Geometry and Semantics t | Video understanding aims to enable models to perceive, reason about, a | Zhaochong An, Zirui Li+ | 5.2 |
| 🔥 | Audio-Visual Flamingo: Open Audio-Visual Intellige | We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of- | Sreyan Ghosh, Arushi Goel+ | 5.2 |
| 🔥 | Sparse Evidence Can Suffice: Agentic Evidence Seek | Multimodal video misinformation detection is commonly formulated as a | Haochen Zhao, Yongxiu Xu+ | 5.2 |
| 🔥 | HAS: Highlight-guided Attention Steering for Multi | Video understanding has become more and more important with the growth | Rui Chu, Yingjie Lao | 5.2 |
| 🔥 | Time-Reversed Imaging: A Multimodal Benchmark and | We introduce time-reversed imaging, a new paradigm that infers what ju | Jorge Bacca, Kebin Contreras+ | 5.2 |
| 🔥 | Test-Time Adaptation via Dual Distillation for Vid | Deep learning models have achieved state-of-the-art performance in sev | André Sacilotti, Samuel Felipe dos Santos+ | 5.2 |
| 🔥 | CADER: Confidence-Aware Dynamic Evidence Reasoning | Long-video understanding increasingly relies on large vision-language | Jinlong Yang, Wenhao Zhang+ | 5.2 |
Showing 12 of 253 papers. See ALL_PAPERS.md for all entries.
📅 2025 (106 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Enhancing Video Understanding: Deep Neural Network | It's no secret that video has become the primary way we share informat | Amir Hosein Fadaei, Mohammad-Reza A. Dehaqani | 5.8 |
| 🔥 | Fine tuning 3D Convolutional Networks for enhanced | The study of Human Activity Recognition (HAR) has attracted considerab | Abir Frad, Hend Basly+ | 5.4 |
| 🔥 | A Hybrid 3D CNNs Transformer Architecture for Vide | Video-Based Human Action Recognition (HAR) remains challenging due to | Engin Seven, Eylem Yücel Demirel | 5.4 |
| 🔥 | Video-CoT: A Comprehensive Dataset for Spatiotempo | Video content comprehension is essential for various applications, ran | Shuyi Zhang, Xiaoshuai Hao+ | 5.2 |
| 🔥 | V-STaR: Benchmarking Video-LLMs on Video Spatio-Te | Human processes video reasoning in a sequential spatio-temporal reason | Zixu Cheng, Jian Hu+ | 5.2 |
| 🔥 | Video deepfake detection using a hybrid CNN-LSTM-T | The proliferation of deepfake technology poses significant challenges | G. Petmezas, Vazgken Vanian+ | 5.0 |
| 🔥 | TinyLLaVA-Video: Towards Smaller LMMs for Video Un | Video behavior recognition and scene understanding are fundamental tas | Xingjian Zhang, Xi Weng+ | 4.9 |
| 🔥 | Harnessing Synthetic Preference Data for Enhancing | While Video Large Language Models (Video-LLMs) have demonstrated remar | Sameep Vani, Shreyas Jena+ | 4.9 |
| 🔥 | AceVFI: A Comprehensive Survey of Advances in Vide | Video Frame Interpolation (VFI) is a core low-level vision task that s | Dahyeon Kye, Changhyun Roh+ | 4.9 |
| 🔥 | How Much 3D Do Video Foundation Models Encode? | Videos are continuous 2D projections of 3D worlds. After training on l | Zixuan Huang, Xiang Li+ | 4.9 |
| 🔥 | A Novel 3D Convolutional Neural Network-Based Deep | Accurate analysis of medical videos remains a major challenge in deep | M. K. Dhar, Mou Deb+ | 4.9 |
| 🔥 | RepAttn3D: Re-parameterizing 3D attention with spa | The technique of structural re-parameterization has been widely adopte | Xiusheng Lu, Lechao Cheng+ | 4.8 |
Showing 12 of 106 papers. See ALL_PAPERS.md for all entries.
📅 2024 (102 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | InternVideo2: Scaling Foundation Models for Multim | We introduce InternVideo2, a new family of video foundation models (Vi | Yi Wang, Kunchang Li+ | 5.2 |
| 🔥 | Various frameworks for integrating image and video | Human action recognition has been identified as an important research | Shaimaa Yosry, Lamiaa A. Elrefaei+ | 4.8 |
| 🔥 | Video-based Exercise Classification and Activated | This paper introduces a simple yet effective strategy for exercise cla | Manvik Pasula, Pramit Saha | 4.7 |
| 🔥 | LongVU: Spatiotemporal Adaptive Compression for Lo | Multimodal Large Language Models (MLLMs) have shown promising progress | Xiaoqian Shen, Yunyang Xiong+ | 4.7 |
| 🔥 | VideoMamba: State Space Model for Efficient Video | Addressing the dual challenges of local redundancy and global dependen | Kunchang Li, Xinhao Li+ | 4.7 |
| 🔥 | Isolated Video-Based Sign Language Recognition Usi | Sign language is a complex language that uses hand gestures, body move | Diksha Kumari, Radhey Shyam Anand | 4.7 |
| 🔥 | Can VLMs be used on videos for action recognition? | Recent advancements have introduced multiple vision-language models (V | Harsh Lunia | 4.4 |
| 🔥 | Prompting Video-Language Foundation Models with Do | Video Question Answering (VideoQA) represents a crucial intersection b | Ting Yu, Kunhao Fu+ | 4.4 |
| 🔥 | Relevance-guided Audio Visual Fusion for Video Sal | Audio data, often synchronized with video frames, plays a crucial role | Li Yu, Xuanzhe Sun+ | 4.4 |
| 🔥 | Automated diagnosis of respiratory diseases from l | An automated computerized approach can aid radiologists in the early d | Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan+ | 4.3 |
| 🔥 | Facial Expression Recognition in Video Using 3D-CN | The focus of research work presented in this paper on improving perfor | Sathisha G, C. K. Subbaraya+ | 4.3 |
| 🔥 | Interpretability in Video-based Human Action Recog | Interpretability plays a vital role in understanding complex deep lear | Jorge Garcia-Torres Fernandez | 4.3 |
Showing 12 of 102 papers. See ALL_PAPERS.md for all entries.
📅 2023 (57 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Hierarchical Spatiotemporal Feature Fusion Network | Current video saliency prediction methods have made great progress rel | Yunzuo Zhang, Tian Zhang+ | 5.7 |
| 🔥 | Video-FocalNets: Spatio-Temporal Focal Modulation | Recent video recognition models utilize Transformer models for long-ra | Syed Talal Wasim, Muhammad Uzair Khattak+ | 5.5 |
| 🔥 | Deep Neural Networks in Video Human Action Recogni | Currently, video behavior recognition is one of the most foundational | Zihan Wang, Yang Yang+ | 5.5 |
| 🔥 | Video Understanding with Large Language Models: A | With the burgeoning growth of online video platforms and the escalatin | Yolo Y. Tang, Jing Bi+ | 5.2 |
| 🔥 | Audio-visual Saliency for Omnidirectional Videos | Visual saliency prediction for omnidirectional videos (ODVs) has shown | Yuxin Zhu, Xilei Zhu+ | 5.2 |
| 🔥 | Understanding Video Transformers for Segmentation: | Video segmentation encompasses a wide range of categories of problem f | Rezaul Karim, Richard P. Wildes | 4.8 |
| 🔥 | VMC: Video Motion Customization using Temporal Att | Text-to-video diffusion models have advanced video generation signific | Hyeonho Jeong, Geon Yeong Park+ | 4.7 |
| 🔥 | A Video Is Worth 4096 Tokens: Verbalize Videos To | Multimedia content, such as advertisements and story videos, exhibit a | Aanisha Bhattacharya, Yaman K Singla+ | 4.6 |
| 🔥 | A dynamic gesture recognition method based on R(2+ | Efficient spatial-temporal feature extraction from input video streams | Yupeng Huo, Jie Shen+ | 4.6 |
| 🔥 | Spatio-Temporal Features based Human Action Recogn | —Recognition of human intention is crucial and challenging due to subt | Saifuddin Saif, E. Wollega+ | 4.5 |
| 🔥 | AMS-Net: Modeling Adaptive Multi-Granularity Spati | Effective spatio-temporal modeling as a core of video representation l | Qilong Wang, Qiyao Hu+ | 4.5 |
| 🔥 | Video Traffic Analysis for Real-Time Emotion Recog | Since the outbreak of the COVID-19 crisis, the transition to remote ed | Ayoub Sassi, W. Jaafar+ | 4.5 |
Showing 12 of 57 papers. See ALL_PAPERS.md for all entries.
📅 2022 (51 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Large-scale Robustness Analysis of Video Action Re | We have seen a great progress in video action recognition in recent ye | Madeline Chantry Schiappa, Naman Biyani+ | 5.5 |
| 🔥 | VRT: A Video Restoration Transformer | Video restoration (e. g. , video super-resolution) aims to restore hig | Jingyun Liang, Jiezhang Cao+ | 5.5 |
| 🔥 | 3D Convolutional with Attention for Action Recogni | Human action recognition is one of the challenging tasks in computer v | Labina Shrestha, Shikha Dubey+ | 5.0 |
| 🔥 | UniFormerV2: Spatiotemporal Learning by Arming Ima | Learning discriminative spatiotemporal representation is the key probl | Kunchang Li, Yali Wang+ | 5.0 |
| 🔥 | Video Human Action Recognition Algorithm Based on | The traditional action recognition algorithm based on manual feature e | Yu Wang, Jiaxi Sun | 4.8 |
| 🔥 | No-Reference Video Quality Assessment Using Multi- | With the constantly growing popularity of video-based services and app | D. Varga | 4.8 |
| 🔥 | Action Recognition Using Action Sequences Optimiza | Effective extraction and representation of action information are crit | Xin Xiong, Weidong Min+ | 4.5 |
| 🔥 | Enhancing Deformable Convolution based Video Frame | This paper presents a new deformable convolution-based video frame int | Duolikun Danier, Fan Zhang+ | 4.4 |
| 🔥 | Skeleton Graph-Neural-Network-Based Human Action R | Human action recognition has been applied in many fields, such as vide | Miao Feng, Jean Meunier | 4.4 |
| 🔥 | Two-stream fusion model using 3D-CNN and 2D-CNN vi | Hand gestures are useful tools for many applications in the human-comp | Debajit Sarma, V. Kavyasree+ | 4.3 |
| 🔥 | Sign Language Recognition Based on R(2+1)D With Sp | Previous work utilized three-dimensional (3-D) convolutional neural ne | Xiangzu Han, Fei Lu+ | 4.3 |
| 🔥 | An Effective Video Transformer With Synchronized S | Convolutional neural networks (CNNs) have come to dominate vision-base | S. Alfasly, C. Chui+ | 4.3 |
Showing 12 of 51 papers. See ALL_PAPERS.md for all entries.
📅 2021 (49 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Spatiotemporal Dilated Convolution with Uncertain | In this paper, we propose a novel SpatioTemporal convolutional Dense N | Yu-Jen Ma, Hong-Han Shuai+ | 6.3 |
| 🔥 | Action Transformer: A Self-Attention Model for Sho | Deep neural networks based purely on attention have been successful ac | Vittorio Mazzia, Simone Angarano+ | 5.7 |
| 🔥 | Temporal-attentive Covariance Pooling Networks for | For video recognition task, a global representation summarizing the wh | Zilin Gao, Qilong Wang+ | 5.7 |
| 🔥 | Is Space-Time Attention All You Need for Video Und | We present a convolution-free approach to video classification built e | Gedas Bertasius, Heng Wang+ | 5.6 |
| 🔥 | Revisiting Video Saliency Prediction in the Deep L | Predicting where people look in static scenes, a. k. a visual saliency | Wenguan Wang, Jianbing Shen+ | 5.3 |
| 🔥 | TAda! Temporally-Adaptive Convolutions for Video U | Spatial convolutions are widely used in numerous deep video models. It | Ziyuan Huang, Shiwei Zhang+ | 5.2 |
| 🔥 | Efficient Action Recognition with Introducing R(2+ | The mainstream methods in video action recognition includes 3D convolu | Hao Jin, Jianming Yang+ | 5.2 |
| 🔥 | Token Shift Transformer for Video Classification | Transformer achieves remarkable successes in understanding 1 and 2-dim | Hao Zhang, Y. Hao+ | 5.0 |
| 🔥 | Improved CNN-based Learning of Interpolation Filte | The versatility of recent machine learning approaches makes them ideal | Luka Murn, Saverio Blasi+ | 4.9 |
| 🔥 | SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-Ki | This report presents the technical details of our submission to the EP | Swathikiran Sudhakaran, Adrian Bulat+ | 4.7 |
| 🔥 | "Knights": First Place Submission for VIPriors21 A | This technical report presents our approach "Knights" to solve the act | Ishan Dave, Naman Biyani+ | 4.4 |
| 🔥 | Recent Advances in Video Action Recognition with 3 | SUMMARY The performance of video action recognition has improved signi | Kensho Hara | 4.3 |
Showing 12 of 49 papers. See ALL_PAPERS.md for all entries.
📅 2020 (41 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Deep Analysis of CNN-based Spatio-temporal Represe | In recent years, a number of approaches based on 2D or 3D convolutiona | Chun-Fu Chen, Rameswar Panda+ | 5.8 |
| 🔥 | Unified Image and Video Saliency Modeling | Visual saliency modeling for images and videos is treated as two indep | Richard Droste, Jianbo Jiao+ | 5.7 |
| 🔥 | TAM: Temporal Adaptive Module for Video Recognitio | Video data is with complex temporal dynamics due to various factors su | Zhaoyang Liu, Limin Wang+ | 5.7 |
| 🔥 | Dissected 3D CNNs: Temporal Skip Connections for E | Convolutional Neural Networks with 3D kernels (3D-CNNs) currently achi | Okan Köpüklü, Stefan Hörmann+ | 5.5 |
| 🔥 | TEA: Temporal Excitation and Aggregation for Actio | Temporal modeling is key for action recognition in videos. It normally | Yan Li, Bin Ji+ | 5.5 |
| 🔥 | RANP: Resource Aware Neuron Pruning at Initializat | Although 3D Convolutional Neural Networks (CNNs) are essential for mos | Zhiwei Xu, Thalaiyasingam Ajanthan+ | 5.2 |
| 🔥 | Developing Motion Code Embedding for Action Recogn | In this work, we propose a motion embedding strategy known as motion c | Maxat Alibayev, David Paulius+ | 5.2 |
| 🔥 | Would Mega-scale Datasets Further Enhance Spatiote | How can we collect and use a video dataset to further improve spatiote | Hirokatsu Kataoka, Tenga Wakamiya+ | 5.0 |
| 🔥 | Learnable Sampling 3D Convolution for Video Enhanc | A key challenge in video enhancement and action recognition is to fuse | Shuyang Gu, Jianmin Bao+ | 5.0 |
| 🔥 | Challenge report:VIPriors Action Recognition Chall | This paper is a brief report to our submission to the VIPriors Action | Zhipeng Luo, Dawei Xu+ | 4.7 |
| 🔥 | Res3ATN -- Deep 3D Residual Attention Network for | Hand gesture recognition is a strenuous task to solve in videos. In th | Naina Dhingra, Andreas Kunz | 4.6 |
| 🔥 | Toward Accurate Person-level Action Recognition in | Detecting and recognizing human action in videos with crowded scenes i | Li Yuan, Yichen Zhou+ | 4.4 |
Showing 12 of 41 papers. See ALL_PAPERS.md for all entries.
📅 2019 (27 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Spatio-Temporal FAST 3D Convolutions for Human Act | Effective processing of video input is essential for the recognition o | Alexandros Stergiou, Ronald Poppe | 5.8 |
| 🔥 | A review of Convolutional-Neural-Network-based act | Abstract Video action recognition is widely applied in video indexing, | Guangle Yao, Tao Lei+ | 5.3 |
| 🔥 | Spatiotemporal distilled dense-connectivity networ | Abstract Two-stream convolutional neural networks show great promise f | Wangli Hao, Zhaoxiang Zhang | 5.3 |
| 🔥 | Resource Efficient 3D Convolutional Neural Network | Recently, convolutional neural networks with 3D kernels (3D CNNs) have | Okan Köpüklü, Neslihan Kose+ | 5.0 |
| 🔥 | Image and Video Compression with Neural Networks: | In recent years, the image and video coding technologies have advanced | Siwei Ma, Xinfeng Zhang+ | 4.9 |
| 🔥 | Improving Action Recognition with the Graph-Neural | Recent human action recognition methods mainly model a two-stream or 3 | Wu Luo, Chongyang Zhang+ | 4.5 |
| 🔥 | Multi-teacher Knowledge Distillation for Compresse | Recently, convolutional neural networks (CNNs) have seen great progres | Meng-Chieh Wu, C. Chiu+ | 4.2 |
| 🔥 | Motion Sickness Prediction in Stereoscopic Videos | In this paper, we propose a three-dimensional (3D) convolutional neura | Tae Min Lee, Jong-Chul Yoon+ | 4.2 |
| 🔥 | Predicting 3D Human Dynamics from Video | Given a video of a person in action, we can easily guess the 3D future | Jason Y. Zhang, Panna Felsen+ | 4.1 |
| 🔥 | FBK-HUPBA Submission to the EPIC-Kitchens 2019 Act | In this report we describe the technical details of our submission to | Swathikiran Sudhakaran, Sergio Escalera+ | 4.1 |
| 📎 | Explainable Deep Learning for Video Recognition Ta | The popularity of Deep Learning for real-world applications is ever-gr | Liam Hiley, Alun Preece+ | 3.9 |
| 📎 | Deep 3D Convolutional Neural Network for Automated | Computer Aided Diagnosis has emerged as an indispensible technique for | Sumita Mishra, Naresh Kumar Chaudhary+ | 3.9 |
Showing 12 of 27 papers. See ALL_PAPERS.md for all entries.
📅 2018 (28 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Interpretable Spatio-temporal Attention for Video | Inspired by the observation that humans are able to process videos eff | Lili Meng, Bo Zhao+ | 7.4 |
| 🔥 | Revisiting Video Saliency: A Large-scale Benchmark | In this work, we contribute to video saliency research in two ways. Fi | Wenguan Wang, Jianbing Shen+ | 5.3 |
| 🔥 | Review of Visual Saliency Detection with Comprehen | Visual saliency detection model simulates the human visual system to p | Runmin Cong, Jianjun Lei+ | 5.0 |
| 🔥 | Reduced-Gate Convolutional LSTM Using Predictive C | Spatiotemporal sequence prediction is an important problem in deep lea | Nelly Elsayed, Anthony S. Maida+ | 4.9 |
| 🔥 | Recurrent Convolutions for Causal 3D CNNs | Recently, three dimensional (3D) convolutional neural networks (CNNs) | Gurkirt Singh, Fabio Cuzzolin | 4.8 |
| 🔥 | Non-local NetVLAD Encoding for Video Classificatio | This paper describes our solution for the 2$^\text{nd}$ YouTube-8M vid | Yongyi Tang, Xing Zhang+ | 4.6 |
| 🔥 | Morph: Flexible Acceleration for 3D CNN-Based Vide | The past several years have seen both an explosion in the use of Convo | Kartik Hegde, R. Agrawal+ | 4.5 |
| 🔥 | ECO: Efficient Convolutional Network for Online Vi | The state of the art in video understanding suffers from two problems: | Mohammadreza Zolfaghari, Kamaljeet Singh+ | 4.4 |
| 🔥 | Non-Local Video Denoising by CNN | Non-local patch based methods were until recently state-of-the-art for | Axel Davy, Thibaud Ehret+ | 4.4 |
| 🔥 | Recurrence to the Rescue: Towards Causal Spatiotem | Recently, three dimensional (3D) convolutional neural networks (CNNs) | Gurkirt Singh, Fabio Cuzzolin | 4.3 |
| 🔥 | Benchmark 3D eye-tracking dataset for visual salie | Visual Attention Models (VAMs) predict the location of an image or vid | Amin Banitalebi-Dehkordi, Eleni Nasiopoulos+ | 4.2 |
| 🔥 | SlowFast Networks for Video Recognition | We present SlowFast networks for video recognition. Our model involves | Christoph Feichtenhofer, Haoqi Fan+ | 4.2 |
Showing 12 of 28 papers. See ALL_PAPERS.md for all entries.
📅 2017 (19 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | The Monkeytyping Solution to the YouTube-8M Video | This article describes the final solution of team monkeytyping, who fi | He-Da Wang, Teng Zhang+ | 5.2 |
| 🔥 | A Closer Look at Spatiotemporal Convolutions for A | In this paper we discuss several forms of spatiotemporal convolutions | Du Tran, Heng Wang+ | 5.1 |
| 🔥 | Predicting Video Saliency with Object-to-Motion CN | Over the past few years, deep neural networks (DNNs) have exhibited gr | Lai Jiang, Mai Xu+ | 5.0 |
| 🔥 | Hierarchical Deep Recurrent Architecture for Video | This paper introduces the system we developed for the Youtube-8M Video | Luming Tang, Boyang Deng+ | 4.9 |
| 🔥 | Graph-Theoretic Spatiotemporal Context Modeling fo | As an important and challenging problem in computer vision, video sali | Lina Wei, Fangfang Wang+ | 4.8 |
| 🔥 | Video Classification With CNNs: Using The Codec As | We investigate video classification via a two-stream convolutional neu | Aaron Chadha, Alhabib Abbas+ | 4.7 |
| 🔥 | Facial Expression Recognition Using Enhanced Deep | Deep Neural Networks (DNNs) have shown to outperform traditional metho | Behzad Hasani, Mohammad H. Mahoor | 4.4 |
| 🔥 | Temporal Relational Reasoning in Videos | Temporal relational reasoning, the ability to link meaningful transfor | Bolei Zhou, A. Andonian+ | 4.2 |
| 🔥 | Two-Stream 3D Convolutional Neural Network for Ske | It remains a challenge to efficiently extract spatialtemporal informat | Hong Liu, Juanhui Tu+ | 4.2 |
| 🔥 | Spatio-Temporal Facial Expression Recognition Usin | Automated Facial Expression Recognition (FER) has been a challenging t | Behzad Hasani, Mohammad H. Mahoor | 4.1 |
| 📎 | A Brief Survey of Deep Reinforcement Learning | Deep reinforcement learning is poised to revolutionise the field of AI | Kai Arulkumaran, Marc Peter Deisenroth+ | 3.9 |
| 📎 | Rethinking Spatiotemporal Feature Learning For Vid | Rethinking Spatiotemporal Feature Learning For Video Understanding | Saining Xie, Chen Sun+ | 3.9 |
Showing 12 of 19 papers. See ALL_PAPERS.md for all entries.
📅 2016 (6 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Convolutional Two-Stream Network Fusion for Video | Recent applications of Convolutional Neural Networks (ConvNets) for hu | Christoph Feichtenhofer, A. Pinz+ | 5.0 |
| 🔥 | Large-Scale Shape Retrieval with Sparse 3D Convolu | In this paper we present results of performance evaluation of S3DCNN - | Alexandr Notchenko, Ermek Kapushev+ | 4.4 |
| 📎 | Deep Learning for Saliency Prediction in Natural V | The purpose of this paper is the detection of salient areas in natural | Souad Chaabouni, Jenny Benois-Pineau+ | 3.4 |
| 📎 | Grad-CAM: Visual Explanations from Deep Networks v | We propose a technique for producing ‘visual explanations’ for decisio | Ramprasaath R. Selvaraju, Abhishek Das+ | 3.3 |
| 📎 | A Deep Learning Approach for Joint Video Frame and | Reinforcement learning is concerned with identifying reward-maximizing | Felix Leibfried, Nate Kushman+ | 2.6 |
| 📎 | Transfer learning with deep networks for saliency | Transfer learning with deep networks for saliency prediction in natura | S. Chaabouni, J. Benois-Pineau+ | 2.6 |
📅 2015 (6 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | C3D: Generic Features for Video Analysis | We propose a simple yet effective approach for spatiotemporal feature | Du Tran, Lubomir Bourdev+ | 5.7 |
| 🔥 | Intra-and-Inter-Constraint-based Video Enhancement | Video enhancement plays an important role in various video application | Yuanzhe Chen, Weiyao Lin+ | 4.1 |
| 📎 | Activity Recognition Using A Combination of Catego | This paper presents a novel approach for automatic recognition of huma | Weiyao Lin, Ming-Ting Sun+ | 3.8 |
| 📎 | A new network-based algorithm for human activity r | In this paper, a new network-transmission-based (NTB) algorithm is pro | Weiyao Lin, Yuanzhe Chen+ | 3.8 |
| 📎 | Group Event Detection with a Varying Number of Gro | This paper presents a novel approach for automatic recognition of grou | Weiyao Lin, Ming-Ting Sun+ | 2.9 |
| 📎 | VoxNet: A 3D Convolutional Neural Network for real | VoxNet: A 3D Convolutional Neural Network for real-time object recogni | D. Maturana, S. Scherer | 2.8 |
📅 2014 (2 papers)
| Tag | Title | Summary | Author | Score |
|---|---|---|---|---|
| 🔥 | Visualizing and Understanding Convolutional Networ | Large Convolutional Network models have recently demonstrated impressi | Matthew D. Zeiler, Rob Fergus | 5.4 |
| 📎 | Deep Inside Convolutional Networks: Visualising Im | This paper addresses the visualisation of image classification models, | Karen Simonyan, Andrea Vedaldi+ | 3.7 |
📄 Full paper list: ALL_PAPERS.md
┌─────────────────────────────────────────────────────────────┐
│ search_config.json │
│ (24 queries × 3 depth layers) │
└───────────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Collector │
│ arXiv API + Semantic Scholar API │
└───────────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Normalizer + CrossRef Enrichment │
│ ID/version normalization, dedup, citations │
└───────────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Scorer │
│ Keyword Match + Citations + Venue + Survey Bonus │
└───────────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Storage │
│ papers/index.jsonl + papers/quarantine.jsonl │
└───────────────────────────┬─────────────────────────────────┘
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│ README │ │ Feishu │ │ Dashboard │
│ (Display)│ │ (Notify) │ │ (HTML) │
└──────────┘ └──────────┘ └──────────┘
| 🎯 Smart Search | 📊 Data Enhancement | 🌐 Multi-Source | 🔔 Auto Notify |
|---|---|---|---|
| Daily arXiv search | CrossRef enrichment for new papers | arXiv + Semantic Scholar | Feishu Webhook |
| 24 layered queries | CN/EN summary generation | Title dedup + ID norm | Success/Failure alerts |
| Score-based filtering | Topic clustering | Citation + Venue boost | GitHub Actions |
GitHub Actions aggregates arXiv and Semantic Scholar daily, then normalizes, deduplicates, scores, and enriches accepted papers with CrossRef.
For academic research use only