EgoSound: Benchmarking Sound Understanding in Egocentric Videos
Zhu, B., Fu, Y., Dong, Q., Sun, G., Qian, T., Wu, Y., Paudel, D. P., Fu, Y., & Xue, X. (2026). EgoSound: Benchmarking Sound Understanding in Egocentric Videos. In CVPR.
Zhu, B., Fu, Y., Dong, Q., Sun, G., Qian, T., Wu, Y., Paudel, D. P., Fu, Y., & Xue, X. (2026). EgoSound: Benchmarking Sound Understanding in Egocentric Videos. In CVPR.
Zhang, D., Fu, Y., Yang, R., Miao, Y., Qian, T., Zheng, X., Sun, G., Chhatkuli, A., Huang, X., Jiang, Y.-G., Van Gool, L., & Paudel, D. P. (2026). EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark. In ICLR.
Zhou, Y., Li, Y., Fu, Y., Benini, L., Konukoglu, E., & Sun, G. (2025). CamSAM2: Segment Anything Accurately in Camouflaged Videos. In NeurIPS.
Tan, Y., Wu, Z., Fu, Y., Zhou, Z., Sun, G., Zamfir, E., Ma, C., Paudel, D. P., Van Gool, L., & Timofte, R. (2025). XTrack: Multimodal Training Boosts RGB-X Video Object Trackers. In ICCV.
Fu, Y., Wang, R., Ren, B., Sun, G., Gong, B., Fu, Y., Paudel, D. P., Huang, X., & Van Gool, L. (2025). ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives. In ICCV.
An, Z., Sun, G., Liu, Y., Li, R., Han, J., Konukoglu, E., & Belongie, S. (2025). Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model. In CVPR.
An, Z., Sun, G., Liu, Y., Li, R., Wu, M., Cheng, M.-M., Konukoglu, E., & Belongie, S. (2025). Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation. In ICLR.
Li, X., Liu, Y., Sun, G., Wu, M., Zhang, L., & Zhu, C. (2025). Towards Open-Vocabulary Video Semantic Segmentation. IEEE Transactions on Multimedia, 27, 2924–2934.
Sun, G., An, Z., Liu, Y., Liu, C., Sakaridis, C., Fan, D. P., & Van Gool, L. (2023). Indiscernible Object Counting in Underwater Scenes. In CVPR.
Tang, H., Sun, G., Sebe, N., & Van Gool, L. (2023). Edge Guided GANs With Multi-Scale Contrastive Learning for Semantic Image Synthesis. TPAMI.
Sun, G., Liu, Y., Tang, H., Chhatkuli, A., Zhang, L., & Van Gool, L. (2022). Mining relations among cross-frame affinities for video semantic segmentation. In ECCV.
Sun, G., Liu, Y., Ding, H., Probst, T., & Van Gool, L. (2022). Coarse-to-fine feature mining for video semantic segmentation. In CVPR.
Wang, W., Sun, G., & Van Gool, L. (2022). Looking beyond single images for weakly supervised semantic segmentation learning. TPAMI.
Liang, J., Cao, J., Sun, G., Zhang, K., Van Gool, L., & Timofte, R. (2021). Swinir: Image restoration using swin transformer. In ICCVW.
Sun, G., Probst, T., Paudel, D. P., Popović, N., Kanakis, M., Patel, J., ... & Van Gool, L. (2021). Task switching network for multi-task learning. In ICCV.
Liang, J., Sun, G., Zhang, K., Van Gool, L., & Timofte, R. (2021). Mutual affine network for spatially variant kernel estimation in blind image super-resolution. In ICCV.
Popovic, N., Paudel, D. P., Probst, T., Sun, G., & Van Gool, L. (2021). Compositetasking: Understanding images by spatial composition of tasks. In CVPR.
Cholakkal, H., Sun, G., Khan, S., Khan, F. S., Shao, L., & Van Gool, L. (2020). Towards partial supervision for generic object counting in natural scenes. TPAMI, 44(3), 1604-1622.
Sun, G., Wang, W., Dai, J., & Van Gool, L. (2020). Mining cross-image semantics for weakly supervised semantic segmentation. In ECCV.
Sun, G., Khan, S., Li, W., Cholakkal, H., Khan, F. S., & Van Gool, L. (2020). Fixing localization errors to improve image classification. In ECCV.
Fan, D. P., Ji, G. P., Sun, G., Cheng, M. M., Shen, J., & Shao, L. (2020). Camouflaged object detection. In CVPR.
Sun, G., Cholakkal, H., Khan, S., Khan, F., & Shao, L. (2020, April). Fine-grained recognition: Accounting for subtle differences between similar classes. In AAAI.
Cholakkal, H., Sun, G., Khan, F. S., & Shao, L. (2019). Object counting and instance segmentation with image-level supervision. In CVPR.
Waqas Zamir, S., Arora, A., Gupta, A., Khan, S., Sun, G., Shahbaz Khan, F., ... & Bai, X. (2019). isaid: A large-scale dataset for instance segmentation in aerial images. In CVPRW.