A System for Large-Scale Image and Video Retrieval on Everyday Scenes

Published in University of Missouri-Columbia, 2022

Recommended citation: Arun George Zachariah - "A System for Large-Scale Image and Video Retrieval on Everyday Scenes." Doctoral dissertation, University of Missouri-Columbia, 2022. /publication/Arun_George_Zachariah_Dissertation.pdf

There has been a growing amount of multimedia data generated on the web today in terms of size and diversity. This has made accurate retrieval from these large and complex collections of data a challenging problem. Motivated by the need for systems that can enable scalable and efficient search, this dissertation proposes QIK (Querying Images Using Contextual Knowledge), a system that leverages deep learning and natural language processing for large-scale image and video retrieval of everyday scenes with common objects.

Download the full dissertation PDF

Bibliography

  1. BV Patel and BB Meshram. Content Based Video Retrieval Systems. arXiv preprint arXiv:1205.1641, 2012.
  2. Aasif Ansari and Muzammil H Mohammed. Content Based Video Retrieval Systems-Methods, Techniques, Trends and Challenges. International Journal of Computer Applications, 112(7):13–22, 2015.
  3. Michael S Lew, Nicu Sebe, and John P Eakins. Challenges of Image and Video Retrieval. In International Conference on Image and Video Retrieval, pages 1–6, 2002.
  4. Daniel Jurafsky and James H. Martin. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Prentice Hall, USA, 2nd edition, 2009.
  5. Hyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, and Bohyung Han. Large-Scale Image Retrieval with Attentive Deep Local Features. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 1–10, 2017.
  6. Albert Gordo, Jon Almazán, Jerome Revaud, and Diane Larlus. Deep Image Retrieval: Learning Global Representations for Image Search. In Proceedings of the European Conference on Computer Vision, pages 241–257, 2016.
  7. Yannis Kalantidis, Clayton Mellina, and Simon Osindero. Cross-Dimensional Weighting for Aggregated Deep Convolutional Features. In Computer Vision - ECCV 2016 Workshops, pages 685–701, 2016.
  8. Amaia Salvador, Xavier Giro-i Nieto, Ferran Marques, and Shin'ichi Satoh. Faster R-CNN Features for Instance Search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016.
  9. Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision, pages 740–755, 2014.
  10. Li Yuan, Tao Wang, Xiaopeng Zhang, Francis EH Tay, Zequn Jie, Wei Liu, and Jiashi Feng. Central Similarity Quantization for Efficient Image and Video Retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3083–3092, 2020.
  11. Giorgos Kordopatis-Zilos, Christos Tzelepis, Symeon Papadopoulos, Ioannis Kompatsiaris, and Ioannis Patras. DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval. arXiv preprint arXiv:2106.13266, 2021.
  12. Jun Xu, Tao Mei, Ting Yao, and Yong Rui. MSR-VTT: A Large Video Description Dataset for Bridging Video and Language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  13. Tim Bray, Jean Paoli, C. M. Sperberg-McQueen, and Eve Maler. Extensible Markup Language (XML) 1.0 Second Edition W3C Recommendation. Technical Report REC-xml-20001006, World Wide Web Consortium, Oct 2000.
  14. Anders Berglund, Scott Boag, Don Chamberlin, Mary F. Fernandez, Michael Kay, Jonathan Robie, and Jerome Simeon. XML Path Language (XPath) 2.0 W3C Working Draft 16. Technical Report WD-xpath20-20020816, World Wide Web Consortium, Aug 2002.
  15. Kuo-Chung Tai. The Tree-to-Tree Correction Problem. Journal of the ACM (JACM), 26(3):422–433, 1979.
  16. Benjamin Paaßen, Claudio Gallicchio, Alessio Micheli, and Barbara Hammer. Tree Edit Distance Learning via Adaptive Symbol Embeddings. In Proceedings of the 35th International Conference on Machine Learning, volume 80, pages 3976–3985, 2018.
  17. Philip Bille. A Survey on Tree Edit Distance and Related Problems. Theoretical Computer Science, 337(1):217–239, 2005.
  18. Kaizhong Zhang and Dennis Shasha. Simple Fast Algorithms for the Editing Distance Between Trees and Related Problems. SIAM Journal on Computing, 18(6):1245–1262, 1989.
  19. Philip N Klein. Computing the Edit-Distance Between Unrooted Ordered Trees. In European Symposium on Algorithms, pages 91–102, 1998.
  20. Erik D Demaine, Shay Mozes, Benjamin Rossman, and Oren Weimann. An Optimal Decomposition Algorithm for Tree Edit Distance. ACM Transactions on Algorithms (TALG), 6:1–19, 2009.
  21. Mateusz Pawlik and Nikolaus Augsten. Efficient Computation of the Tree Edit Distance. ACM Transactions on Database Systems, 40(1):3:1–3:40, Mar 2015.
  22. Mateusz Pawlik and Nikolaus Augsten. Tree Edit Distance: Robust and Memory-Efficient. Information Systems, 56:157–173, 2016.
  23. Stefan Schwarz, Mateusz Pawlik, and Nikolaus Augsten. A New Perspective on the Tree Edit Distance. In Similarity Search and Applications, pages 156–170, 2017.
  24. Burton H. Bloom. Space/Time Trade-Offs in Hash Coding with Allowable Errors. Communications of the ACM, 13(7):422–426, 1970.
  25. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1, page 1097–1105, 2012.
  26. Karen Simonyan and Andrew Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. In Proceedings of the 3rd International Conference on Learning Representations, 2015.
  27. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  28. Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient Estimation of Word Representations in Vector Space. arXiv preprint arXiv:1301.3781, 2013.
  29. Sepp Hochreiter and Jürgen Schmidhuber. Long Short-Term Memory. Neural Computation, 9(8):1735–1780, 1997.
  30. Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and Tell: A Neural Image Caption Generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2015.
  31. Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and Tell: Lessons Learned from the 2015 MS COCO Image Captioning Challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4):652–663, 2017.
  32. Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going Deeper With Convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  33. Ron Mokady, Amir Hertz, and Amit H Bermano. Clipcap: Clip Prefix for Image Captioning. arXiv preprint arXiv:2111.09734, 2021.
  34. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models from Natural Language Supervision. arXiv preprint arXiv:2103.00020, 2021.
  35. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language Models are Unsupervised Multitask Learners. OpenAI blog, 1(8):9, 2019.
  36. David G. Lowe. Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision, 60:91–110, 2004.
  37. Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool. SURF: Speeded-Up Robust Features. Computer Vision and Image Understanding, 110(3):346–359, 2008.
  38. Gabriella Csurka, Christopher R. Dance, Lixin Fan, Jutta Willamowski, and Cédric Bray. Visual Categorization with Bags of Keypoints. In In Workshop on Statistical Learning in Computer Vision, ECCV, pages 1–22, 2004.
  39. James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman. Object Retrieval with Large Vocabularies and Fast Spatial Matching. In 2007 IEEE Conference on Computer Vision and Pattern Recognition, pages 1–8, 2007.
  40. James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman. Lost in Quantization: Improving Particular Object Retrieval in Large Scale Image Databases. In 2008 IEEE Conference on Computer Vision and Pattern Recognition, pages 1–8, 2008.
  41. Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the Inception Architecture for Computer Vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  42. Artem Babenko, Anton Slesarev, Alexandr Chigorin, and Victor Lempitsky. Neural Codes for Image Retrieval. In Proceedings of the European Conference on Computer Vision, pages 584–599, 2014.
  43. Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  44. Artem Babenko Yandex and Victor Lempitsky. Aggregating Local Deep Features for Image Retrieval. In Proceedings of the IEEE International Conference on Computer Vision, 2015.
  45. Filip Radenovic, Giorgos Tolias, and Ondrej Chum. Fine-Tuning CNN Image Retrieval with No Human Annotation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(7):1655–1668, 2019.
  46. Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez. Aggregating Local Descriptors into a Compact Image Representation. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 3304–3311, 2010.
  47. Relja Arandjelovic and Andrew Zisserman. All About VLAD. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2013.
  48. Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. NetVLAD: CNN Architecture for Weakly Supervised Place Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  49. Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards Real-time Object Detection with Region Proposal Networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems, pages 91–99, 2015.
  50. Marvin Teichmann, Andre Araujo, Menglong Zhu, and Jack Sim. Detect-To-Retrieve: Efficient Regional Aggregation for Image Search. In The IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  51. Giorgos Tolias, Yannis Avrithis, and Herve Jegou. Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images. International Journal of Computer Vision, 116(3):247–261, 2016.
  52. Anastasiya Mishchuk, Dmytro Mishkin, Filip Radenovic, and Jiři Matas. Working Hard to Know Your Neighbor's Margins: Local Descriptor Learning Loss. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 4829–4840, 2017.
  53. Dmytro Mishkin, Filip Radenović, and Jiři" Matas. Repeatability Is Not Enough: Learning Affine Regions via Discriminability. In Proceedings of the European Conference on Computer Vision, pages 287–304, 2018.
  54. Ondrej Chum, James Philbin, Josef Sivic, Michael Isard, and Andrew Zisserman. Total Recall: Automatic Query Expansion with a Generative Feature Model for Object Retrieval. In 2007 IEEE 11th International Conference on Computer Vision, pages 1–8, 2007.
  55. Ondřej Chum, Andrej Mikulík, Michal Perdoch, and Jiří Matas. Total Recall II: Query Expansion Revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 889–896, 2011.
  56. Giorgos Tolias, Ronan Sicre, and Herve Jegou. Particular Object Retrieval with Integral Max-Pooling of CNN Activations. In International Conference on Learning Representations, pages 1–12, 2016.
  57. Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. In Proceedings of the 32nd International Conference on Machine Learning, pages 2048–2057, 2015.
  58. Albert Gordo and Diane Larlus. Beyond Instance-Level Image Retrieval: Leveraging Captions to Learn a Global Visual Representation for Semantic Retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  59. Quoc-Tuan Truong and Lauw Hady W. VistaNet: Visual Aspect Attention Network for Multimodal Sentiment Analysis. In The Thirty-third AAAI Conference, pages 305–312, 2019.
  60. Micah Hodosh, Peter Young, and Julia Hockenmaier. Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics. Journal of Artificial Intelligence Research, 47:853–899, 2013.
  61. Richard Socher, Andrej Karpathy, Quoc V Le, Christopher D Manning, and Andrew Y Ng. Grounded Compositional Semantics for Finding and Describing Images with Sentences. Transactions of the Association for Computational Linguistics, 2:207–218, 2014.
  62. Han Zhu, Mingsheng Long, Jianmin Wang, and Yue Cao. Deep Hashing Network for Efficient Similarity Retrieval. Proceedings of the AAAI Conference on Artificial Intelligence, 30(1), 2016.
  63. Mathias Lux and Savvas A. Chatzichristofis. LIRe: Lucene Image Retrieval: An Extensible Java CBIR Library. In Proceedings of the 16th ACM International Conference on Multimedia, pages 1085–1088, 2008.
  64. Jing Huang, S Ravi Kumar, Mandar Mitra, Wei-Jing Zhu, and Ramin Zabih. Image Indexing using Color Correlograms. In Proceedings of IEEE computer society conference on Computer Vision and Pattern Recognition, pages 762–768. IEEE, 1997.
  65. Shih-Fu Chang, Thomas Sikora, and Atul Purl. Overview of the MPEG-7 Standard. IEEE Transactions on Circuits and Systems for Video Technology, 11(6):688–695, 2001.
  66. Hideyuki Tamura, Shunji Mori, and Takashi Yamawaki. Textural Features Corresponding to Visual Perception. IEEE Transactions on Systems, Man, and Cybernetics, 8(6):460–473, 1978.
  67. Savvas A Chatzichristofis and Yiannis S Boutalis. CEDD: Color and Edge Directivity Descriptor: A Compact Descriptor for Image Indexing and Retrieval. In International Conference on Computer Vision Systems, pages 312–322. Springer, 2008.
  68. Xiao Wu, Alexander G Hauptmann, and Chong-Wah Ngo. Practical Elimination of Near-Duplicates from Web Video Search. In Proceedings of the 15th ACM International Conference on Multimedia, pages 218–227, 2007.
  69. Zi Huang, Heng Tao Shen, Jie Shao, Bin Cui, and Xiaofang Zhou. Practical Online Near-Duplicate Subsequence Detection for Continuous Video Streams. IEEE Transactions on Multimedia, 12(5):386–398, 2010.
  70. Jérôme Revaud, Matthijs Douze, Cordelia Schmid, and Hervé Jégou. Event Retrieval in Large Video Collections with Circulant Temporal Encoding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2459–2466, 2013.
  71. Zhanning Gao, Gang Hua, Dongqing Zhang, Nebojsa Jojic, Le Wang, Jianru Xue, and Nanning Zheng. ER3: A Unified Framework for Event Retrieval, Recognition and Recounting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2253–2262, 2017.
  72. Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, and Yiannis Kompatsiaris. Near-Duplicate Video Retrieval by Aggregating Intermediate CNN Layers. In International Conference on Multimedia Modeling, pages 251–263, 2017.
  73. Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, and Yiannis Kompatsiaris. Near-Duplicate Video Retrieval with Deep Metric Learning. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 347–356, 2017.
  74. Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, and Ioannis Kompatsiaris. ViSiL: Fine-Grained Spatio-Temporal Video Similarity Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6351–6360, 2019.
  75. Tengda Han, Weidi Xie, and Andrew Zisserman. Video Representation Learning by Dense Predictive Coding. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Oct 2019.
  76. Haofei Kuang, Yi Zhu, Zhi Zhang, Xinyu Li, Joseph Tighe, Soren Schwertfeger, Cyrill Stachniss, and Mu Li. Video Contrastive Learning with Global Context. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3195–3204, 2021.
  77. Jingkuan Song, Yi Yang, Zi Huang, Heng Tao Shen, and Richang Hong. Multiple Feature Hashing for Real-Time Large Scale Near-Duplicate Video Retrieval. In Proceedings of the 19th ACM International Conference on Multimedia, pages 423–432, 2011.
  78. Liangliang Cao, Zhenguo Li, Yadong Mu, and Shih-Fu Chang. Submodular Video Hashing: A Unified Framework Towards Video Pooling and Indexing. In Proceedings of the 20th ACM International Conference on Multimedia, pages 299–308, 2012.
  79. Guangnan Ye, Dong Liu, Jun Wang, and Shih-Fu Chang. Large-Scale Video Hashing via Structure Learning. In Proceedings of the IEEE International Conference on Computer Vision, pages 2272–2279, 2013.
  80. Venice Erin Liong, Jiwen Lu, Yap-Peng Tan, and Jie Zhou. Deep Video Hashing. IEEE Transactions on Multimedia, 19(6):1209–1219, 2017.
  81. Yun Gu, Chao Ma, and Jie Yang. Supervised Recurrent Hashing for Large Scale Video Retrieval. In Proceedings of the 24th ACM International Conference on Multimedia, pages 272–276, 2016.
  82. Naifan Zhuang, Jun Ye, and Kien A Hua. DLSTM Approach to Video Modeling with Hashing for Large-Scale Video Retrieval. In 2016 23rd International Conference on Pattern Recognition (ICPR), pages 3222–3227, 2016.
  83. Vivek Veeriah, Naifan Zhuang, and Guo-Jun Qi. Differential Recurrent Neural Networks for Action Recognition. In Proceedings of the IEEE International Conference on Computer Vision, pages 4041–4049, 2015.
  84. Jie Qin, Li Liu, Mengyang Yu, Yunhong Wang, and Ling Shao. Fast Action Retrieval from Videos via Feature Disaggregation. Computer Vision and Image Understanding, 156:104–116, 2017.
  85. Hanwang Zhang, Meng Wang, Richang Hong, and Tat-Seng Chua. Play and Rewind: Optimizing Binary Representations of Videos by Self-Supervised Temporal Hashing. In Proceedings of the 24th ACM International Conference on Multimedia, pages 781–790, 2016.
  86. Shuyan Li, Zhixiang Chen, Jiwen Lu, Xiu Li, and Jie Zhou. Neighborhood Preserving Hashing for Scalable Video Retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8212–8221, 2019.
  87. Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov. Unsupervised Learning of Video Representations using LSTMs. In Proceedings of the 32nd International Conference on Machine Learning, volume 37, pages 843–852, 2015.
  88. Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo. Spatio-temporal Video Re-localization by Warp LSTM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1288–1297, 2019.
  89. Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng. Multi-Fiber Networks for Video Recognition. In Proceedings of the European Conference on Computer Vision, 2018.
  90. Herve Jegou, Matthijs Douze, and Cordelia Schmid. Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search. In Proceedings of the 10th European Conference on Computer Vision: Part I, pages 304–317, 2008.
  91. Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  92. Richard Socher, John Bauer, Christopher D. Manning, and Andrew Y. Ng. Parsing with Compositional Vector Grammars. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, pages 455–465, 2013.
  93. Danqi Chen and Christopher Manning. A Fast and Accurate Dependency Parser using Neural Networks. In Proc. of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 740–750, 2014.
  94. Christian Grün, Sebastian Gath, Alexander Holupirek, and Marc H. Scholl. XQuery Full Text Implementation in BaseX. In Proceedings of the 6th International XML Database Symposium on Database and XML Technologies, pages 114–128, 2009.
  95. Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018.
  96. Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. Learning Transferable Architectures for Scalable Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2018.
  97. Edwin G. Ng, Bo Pang, Piyush Sharma, and Radu Soricut. Understanding Guided Image Captioning Performance across Domains. arXiv preprint arXiv:2012.02339, 2020.
  98. Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. The Design and Operation of CloudLab. In 2019 USENIX Annual Technical Conference (USENIX ATC 19), pages 1–14, 2019.
  99. Steven M. Beitzel, Eric C. Jensen, and Ophir Frieder. MAP, pages 1691–1692. Springer US, Boston, MA, 2009.
  100. Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. Universal Sentence Encoder for English. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 169–174, 2018.
  101. Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, and Tatsuya Harada. MFNet: Towards Real-Time Semantic Segmentation for Autonomous Vehicles with Multi-Spectral Scenes. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5108–5115, 2017.
  102. Arun Zachariah, Praveen Rao, Anas Katib, Monica Senapati, and Kobus Barnard. A Gossip-Based System for Fast Approximate Score Computation in Multinomial Bayesian Networks. In Proceedings of the 35th IEEE International Conference on Data Engineering (ICDE), 2019.
  103. Mohamed Gharibi, Arun Zachariah, and Praveen Rao. FoodKG: A Tool to Enrich Knowledge Graphs Using Machine Learning Techniques. Frontiers in Big Data, 3, 2020.
  104. Suveen Angraal, Arun George Zachariah, Raaisa Raaisa, Rohan Khera, Praveen Rao, Harlan M. Krumholz, and John A. Spertus. Evaluation of Internet-Based Crowdsourced Fundraising to Cover Health Care Costs in the United States. JAMA Network Open, 4(1):e2033157–e2033157, 2021.
  105. Arun Zachariah and Maha Alrasheed. Private-Share: A Secure and Privacy-Preserving De-Centralized Framework for Large Scale Data Sharing. In Proceedings of the 3rd ACM International Conference on Multimedia in Asia, 2021.
  106. Arun Zachariah and Praveen Rao. Large-Scale Image and Video Retrieval on Everyday Scenes With Common Objects. ACM Transactions on Intelligent Systems and Technology, 2022. Under Review.
  107. Arun Zachariah, Mohamed Gharibi, and Praveen Rao. A Large-Scale Image Retrieval System for Everyday Scenes. In Proceedings of the 2nd ACM International Conference on Multimedia in Asia, 2021.
  108. Arun Zachariah, Mohamed Gharibi, and Praveen Rao. QIK: A System for Large-Scale Image Retrieval on Everyday Scenes With Common Objects. In Proceedings of the 2020 International Conference on Multimedia Retrieval, page 126–135, 2020.
  109. Arun Zachariah, Praveen Rao, Brian Corn, and Dominique Davison. Zero Shot Learning for Predicting Energy Usage of Buildings in Sustainable Design. arXiv preprint arXiv:2202.05206, 2022.
  110. Praveen Rao, Arun Zachariah, Deepthi Rao, Peter Tonellato, Wesley Warren, and Eduardo Simoes. Accelerating Variant Calling on Human Genomes Using a Commodity Cluster. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management (CIKM), page 3388–3392, 2021.
  111. Praveen Rao and Arun Zachariah. Enabling Large-Scale Human Genome Sequence Analysis on CloudLab. In IEEE INFOCOM 2022 - IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS), pages 1–2, 2022.
  112. Daniel E. Lopez Barron, Praveen Rao, Deepthi Rao, Ossama Tawfik, and Arun Zachariah. Large-Scale Storage of Whole Slide Images and Fast Retrieval of Tiles using DRAM. In Big Data II: Learning, Analytics, and Applications, pages 45 – 50, 2020.
  113. Nouf Alrasheed, Arun Zachariah, Shivika Prasanna, Deepthi Rao, and Praveen Rao. Deepfakes for Histopathology Images: Myth or Reality? In Proceedings of the IEEE Applied Imagery Pattern Recognition Workshop (AIPR), pages 1 – 7, 2020.