Abstract
Recent progress in artificial intelligence (AI) tools and systems has been significant, especially in their reasoning and efficiency. Notable examples include generative AI-based large language models (LLMs) like Generative Pre-trained Transformer 3.5 (GPT-3.5), GPT-4, and Gemini, among others. In our work, we evaluated the effectiveness of fine-tuned deep learning models compared to general-purpose LLMs in moderating image-based content. We used deep learning models such as convolutional neural networks, ResNet50, and VGG-16, trained them for violence detection on an image dataset, and tested them on a separate dataset. The same test dataset was also evaluated using Large Language and Vision Assistant (LLaVa) and GPT-4, two LLMs that can process images. The results demonstrate that VGG-16 model had the highest accuracy at 0.94, while LLaVa had the lowest at 0.66. GPT-4 showed superiority over LLaVa with an accuracy value of 0.9242. LLaVa recorded the highest precision of all models.
| Original language | English |
|---|---|
| Pages (from-to) | 70-80 |
| Number of pages | 11 |
| Journal | IEEE Intelligent Systems |
| Volume | 39 |
| Issue number | 6 |
| DOIs | |
| State | Published - 2024 |
| Externally published | Yes |
Bibliographical note
Publisher Copyright:© 2001-2011 IEEE.
ASJC Scopus subject areas
- Computer Networks and Communications
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Are Foundation Models the Next-Generation Social Media Content Moderators?'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver