Pragmatic Analysis of AI Image-to-Text Generation: The Case of ChatGP
DOI:
https://doi.org/10.70922/x2sdfx87Keywords:
generative artificial intelligence, I2T generation, pragmatic analysis, inferentiality, referentialityAbstract
The study aims to conduct a comparative analysis of ChatGPT, Grok, and Gemini to evaluate their capacities in image-to-text generation. The researcher utilized three AI-generated image prompts from Canva, which were processed through the mentioned generative AI (GenAI) models. With the aid of a standardized text prompt, each model was instructed to produce eighteen descriptive sentences corresponding to the prompts. The study examined the contextual alignment, pragmatic appropriateness, as well hallucinations and pragmatic errors present in the generated sentences through pragmatic analysis. The analysis revealed the use of appropriate referents, modal markers, and descriptive details that contributed to contextually and pragmatically accurate descriptions of the images. Despite this, several structural inconsistencies and lexical ambiguities surfaced. While the outputs demonstrated logical coherence through the recognition of lexical patterns, the findings indicate that the models heavily rely on probabilistic associations rather than genuine semantic understanding. The limitations of GenAI became evident in descriptions that contained hallucinations, excessive inference, and syntactic errors, particularly in cases where visual evidence was absent or insufficient to support the inferred content. Overall, Gemini outperformed the other models in producing accurate and meaningful descriptions with high contextual and pragmatic precision, followed by ChatGPT, whose outputs showed some structural flaws, Grok produced the highest number of hallucinated outputs and pragmatic errors.
Downloads
References
Alikhani, M., Khalid, B., & Stone, M. (2023). Image-text coherence and its implications for multimodal AI. Frontiers in Artificial Intelligence, 6, 1-13. https://doi.org/10.3389/frai.2023.1048874
Almazán, J. (2019). Evidentiality in Tagalog. [Doktral na Disertasyon, Universidad Autónoma de Madrid]. AUM Repositorio.
Alonso, I., Salabierra, A., Azkune, G., Barnes, J., & de Lacalle, O. (2025). Vision-language models struggle to align entities across modalities. ArXiv. http://dx.doi.org/10.48550/arXiv.2503.03854
Andrada, M. (2021). Subersibong potensiyal ng makina ng pagsasalin: Google translate at tula ni Carlos Bulosan. Malay Journal, 33(2), 29-45. https://doi.org/10.59588/2243-7851.1071
Assadian, B. (2021). Deflationary reference and referential indeterminacy. Sa M. Mojtahedi, S. Rahman, & M. Zarepour (mga Pat.), Mathematics, Logic, and their Philosophies. Logic, Epistemology, and the Unity of Science, 49 (pp. 365-377). Springer, Cham. https://doi.org/10.1007/978-3-030-53654-1_13
Bergelson, E., & Aslin, R. (2017). Semantic specificity in one-year-old’s word comprehension. Language Learning and Development: The Official Journal of the Society for Language Development, 13(4), 481-501. https://doi.org/10.1080/15475441.2017.1324308
Bosque, A. (2024). Analisis sa lexical choice ng ChatGPT-generated na pagsasalin. Malay Journal, 37(1): 94-106. https://doi.org/10.59588/2243-7851.1008
Briana, J. (2024). Is ChatGPT-produced text authentic? a contrastive analysis of cohesive markers in human and AI-generated text. Journal of English and Applied Linguistics, 3(2): 38-54. https://doi.org/10.59588/2961-3094.1120
Chen W., Hu H., Li Y., et al. (2023). Subject-driven text-to-image generation via apprenticeship learning. arXiv. https://doi.org/10.48550/arXiv.2304.00186
Corpus, S. and Villanueva, A. (2024). Speech emotion recognition in Filipino spoken language using deep learning. Sa 2024 IEEE International Conference on Artificial Intelligence in Engineering and Technology, 430-435. https://doi.org/10.1109/IICAIET62352.2024.10730717
Csibra, G., & Gergely, G. (2013). Teleological understanding of actions. Sa M. Banaji & S. Gelman (mga Pat.), Navigating the Social World: What Infants, Children and Other Species Can Teach Us (pp. 38-43). Oxford Academic. https://doi.org/10.1093/acprof:oso/9780199890712.003.0008
Duffy, G. (2008). Pragmatic analysis. Sa A. Klotz, & D. Prakash (mga Pat). Qualitative Methods in International Relations (pp. 168-186). London, Palgrave Macmillan.
Endriga, D., & Rosario F. (2023). Gender bias in machine translation: the case of Filipino-English translation in Google Translate. Sa R. Moratto & M. Bacolod (mga Pat). Translation Studies in the Philippines: Navigating a Multilingual Archipelago (pp. 83-99). Routledge.
Estrella, J. (2023). Does artificial intelligence genuinely capture the essence of language? UP Working Papers in Linguistics, 2(1), 120-124. https://linguistics.upd.edu.ph/wp-content/uploads/2023/10/21-Artificial-Intelligence.pdf
Evkaya, O. & de Carvalho, M (2024). Decoding AI: The story of data analysis in ChatGPT. arXiv. https://doi.org/10.48550/arXiv.2304.00186
He Y., Robey A., Murata N., et al. (2024). Automated black-box prompt engineering for personalized text-to-image generation. arXiv. https://doi.org/10.48550/arXiv.2403.19103
Hinzen, W. (2001). The pragmatics of inferential content. Synthese, 128(1-2), 157-181. http://dx.doi.org/10.1023/A:1010362521497
Ji, Z., Li, N., Frieske, R. et al. (2024). Survey of hallucination in natural language generation. ArXiv. https://doi.org/10.48550/arXiv.2202.03629
Lim, Y., & Shim, H. (2024). Addressing image hallucination in text-to-image generation through factual image retrieval. ArXiv. https://doi.org/10.48550/arXiv.2407.10683
Liwanag L., Liwanag G., & Liwanag L.. (2024). AI in anthem: A comparative analysis of the English and Filipino ChatGPT 4 translations from the existing translations of the Philippine National Anthem. Recoletos Multidisciplinary Research Journal, 12(2), 91-102. https://doi.org/10.32871/rmrj2412.02.07
Longpre, S., Perisetla, K. Chen, A., Ramesh, N., DuBois, C., Singh, S. (2021). Entity-based knowledge conflicts in question answering. ArXiv. https://doi.org/10.48550/arXiv.2109.05052
Malik, M., & Isik, L. (2023). Relational visual representations underlie human social interaction recognition. Nature Communications, 14(7317), 1–11. https://doi.org/10.1038/s41467-023-43156-8
Mignot, E. (2012). The conceptualization of natural gender in English. Anglophonia: French Journal of English Studies, 16(32), 39-61. https://doi.org/10.4000/anglophonia.140
Mojapelo, M. (2023). Aspects of referent tracking in Northern Sotho. Sa B. Achiri-Taboh (Pat.), The Bantu Noun Phrase: Issues and Perspectives (pp. 210-226). London: Routledge. http://dx.doi.org/10.4324/9781003254188-13
Montalan, J., Layacan, J., Africa, D., et al. (2025). Batayan: A Filipino NLP benchmark for evaluating large language models arXiv. https://arxiv.org/abs/2502.14911
Nagaya, N. (2007). Information structure and constituent order in Tagalog. Languge and Linguistics, 8(1), 343-372.
Nordquist, R. (2019, Hulyo 3). Definitions and examples of hypernyms in English. ThoughtCo. https://www.thoughtco.com/hypernym-words-term-1690943
Nordquist, R. (2024, Abril 30). What are hyponyms in English? ThoughtCo. https://www.thoughtco.com/hyponym-words-term-1690946
Nordquist, R. (2025, Mayo 11). Definitions and examples of postmodifiers in English grammar. ThoughtCo. https://www.thoughtco.com/postmodifier-grammar-1691519
Oppenlaender, J. (2022). The creativity of text-to-image generation. In Proceedings of the 25th International Academic Mindtrek Conference (Academic Mindtrek ’22) (p. 11). ACM. https://doi.org/10.1145/3569219.3569352
Oppenlaender, J., Linder, R., & Silvennoinen, J. (2024). Prompting AI art: An investigation into the creative skill of prompt engineering. International Journal of Human–Computer Interaction, 1–23. https://doi.org/10.1080/10447318.2024.2431761
Reid, L., & Liao, H. (2004). A brief syntactic typology of Philippine languages. Language and Linguistics, 5(2), 433-490. http://hdl.handle.net/10125/32991
Souza, M. E., & Weigang, L. (2025). Grok, Gemini, ChatGPT, and Deepseek: Comparative and applications in conversational artificial intelligence. Inteligencia Artificial, 2(1), 1–7. http://dx.doi.org/10.5281/zenodo.14885243
Sperber, D., & Wilson, D. (2004). Relevance theory. Sa L. Horn & G. Ward (mga Pat.), Blackwell’s Handbook of Pragmatics (pp. 607-632). Blackwell.
Trojovsky, P., & Naveen, P. (2024). Overview and challenges of machine translation for contextually appropriate translations. iScience, 27(10), 1–25. https://doi.org/10.1016/j.isci.2024.110878
Villemaire, S. (2025). Canva. https://www.canva.com/learn/how-to-convert-text-images-ai-magic/.
Wang, Y., & Zhang, G. (2025). Lightweight text-to-image generation model based on contrastive language-image pre-training embeddings and conditional variational autoencoders. Electronics, 14(11), 1–31. https://doi.org/10.3390/electronics14112185
Warr, M. (2024). Beat bias? Personalization, bias, and generative AI. Sa J. Cohen & G. Solano (Eds.), Proceedings of Society for Information Technology and Teacher Education International Conference (pp. 1481–1488).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ariel Bosque (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Articles published in the SOCIAL SCIENCES AND DEVELOPMENT REVIEW will be Open-Access articles distributed under the terms and conditions of the Creative Commons Attribution-Noncommercial 4.0 International (CC BY-NC 4.0). This allows for immediate free access to the work and permits any user to read, download, copy, distribute, print, search, or link to the full texts of articles, crawl them for indexing, pass them as data to software, or use them for any other lawful purpose.