Pragmatic Analysis of AI Image-to-Text Generation: The Case of ChatGP

Authors

DOI:

https://doi.org/10.70922/x2sdfx87

Keywords:

generative artificial intelligence, I2T generation, pragmatic analysis, inferentiality, referentiality

Abstract

The study aims to conduct a comparative analysis of ChatGPT, Grok, and Gemini to evaluate their capacities in image-to-text generation. The researcher utilized three AI-generated image prompts from Canva, which were processed through the mentioned generative AI (GenAI) models. With the aid of a standardized text prompt, each model was instructed to produce eighteen descriptive sentences corresponding to the prompts. The study examined the contextual alignment, pragmatic appropriateness, as well hallucinations and pragmatic errors present in the generated sentences through pragmatic analysis. The analysis revealed the use of appropriate referents, modal markers, and descriptive details that contributed to contextually and pragmatically accurate descriptions of the images. Despite this, several structural inconsistencies and lexical ambiguities surfaced. While the outputs demonstrated logical coherence through the recognition of lexical patterns, the findings indicate that the models heavily rely on probabilistic associations rather than genuine semantic understanding. The limitations of GenAI became evident in descriptions that contained hallucinations, excessive inference, and syntactic errors, particularly in cases where visual evidence was absent or insufficient to support the inferred content. Overall, Gemini outperformed the other models in producing accurate and meaningful descriptions with high contextual and pragmatic precision, followed by ChatGPT, whose outputs showed some structural flaws, Grok produced the highest number of hallucinated outputs and pragmatic errors.

Downloads

Download data is not yet available.

References

Alikhani, M., Khalid, B., & Stone, M. (2023). Image-text coherence and its implications for multimodal AI. Frontiers in Artificial Intelligence, 6, 1-13. https://doi.org/10.3389/frai.2023.1048874

Almazán, J. (2019). Evidentiality in Tagalog. [Doktral na Disertasyon, Universidad Autónoma de Madrid]. AUM Repositorio.

Alonso, I., Salabierra, A., Azkune, G., Barnes, J., & de Lacalle, O. (2025). Vision-language models struggle to align entities across modalities. ArXiv. http://dx.doi.org/10.48550/arXiv.2503.03854

Andrada, M. (2021). Subersibong potensiyal ng makina ng pagsasalin: Google translate at tula ni Carlos Bulosan. Malay Journal, 33(2), 29-45. https://doi.org/10.59588/2243-7851.1071

Assadian, B. (2021). Deflationary reference and referential indeterminacy. Sa M. Mojtahedi, S. Rahman, & M. Zarepour (mga Pat.), Mathematics, Logic, and their Philosophies. Logic, Epistemology, and the Unity of Science, 49 (pp. 365-377). Springer, Cham. https://doi.org/10.1007/978-3-030-53654-1_13

Bergelson, E., & Aslin, R. (2017). Semantic specificity in one-year-old’s word comprehension. Language Learning and Development: The Official Journal of the Society for Language Development, 13(4), 481-501. https://doi.org/10.1080/15475441.2017.1324308

Bosque, A. (2024). Analisis sa lexical choice ng ChatGPT-generated na pagsasalin. Malay Journal, 37(1): 94-106. https://doi.org/10.59588/2243-7851.1008

Briana, J. (2024). Is ChatGPT-produced text authentic? a contrastive analysis of cohesive markers in human and AI-generated text. Journal of English and Applied Linguistics, 3(2): 38-54. https://doi.org/10.59588/2961-3094.1120

Chen W., Hu H., Li Y., et al. (2023). Subject-driven text-to-image generation via apprenticeship learning. arXiv. https://doi.org/10.48550/arXiv.2304.00186

Corpus, S. and Villanueva, A. (2024). Speech emotion recognition in Filipino spoken language using deep learning. Sa 2024 IEEE International Conference on Artificial Intelligence in Engineering and Technology, 430-435. https://doi.org/10.1109/IICAIET62352.2024.10730717

Csibra, G., & Gergely, G. (2013). Teleological understanding of actions. Sa M. Banaji & S. Gelman (mga Pat.), Navigating the Social World: What Infants, Children and Other Species Can Teach Us (pp. 38-43). Oxford Academic. https://doi.org/10.1093/acprof:oso/9780199890712.003.0008

Duffy, G. (2008). Pragmatic analysis. Sa A. Klotz, & D. Prakash (mga Pat). Qualitative Methods in International Relations (pp. 168-186). London, Palgrave Macmillan.

Endriga, D., & Rosario F. (2023). Gender bias in machine translation: the case of Filipino-English translation in Google Translate. Sa R. Moratto & M. Bacolod (mga Pat). Translation Studies in the Philippines: Navigating a Multilingual Archipelago (pp. 83-99). Routledge.

Estrella, J. (2023). Does artificial intelligence genuinely capture the essence of language? UP Working Papers in Linguistics, 2(1), 120-124. https://linguistics.upd.edu.ph/wp-content/uploads/2023/10/21-Artificial-Intelligence.pdf

Evkaya, O. & de Carvalho, M (2024). Decoding AI: The story of data analysis in ChatGPT. arXiv. https://doi.org/10.48550/arXiv.2304.00186

He Y., Robey A., Murata N., et al. (2024). Automated black-box prompt engineering for personalized text-to-image generation. arXiv. https://doi.org/10.48550/arXiv.2403.19103

Hinzen, W. (2001). The pragmatics of inferential content. Synthese, 128(1-2), 157-181. http://dx.doi.org/10.1023/A:1010362521497

Ji, Z., Li, N., Frieske, R. et al. (2024). Survey of hallucination in natural language generation. ArXiv. https://doi.org/10.48550/arXiv.2202.03629

Lim, Y., & Shim, H. (2024). Addressing image hallucination in text-to-image generation through factual image retrieval. ArXiv. https://doi.org/10.48550/arXiv.2407.10683

Liwanag L., Liwanag G., & Liwanag L.. (2024). AI in anthem: A comparative analysis of the English and Filipino ChatGPT 4 translations from the existing translations of the Philippine National Anthem. Recoletos Multidisciplinary Research Journal, 12(2), 91-102. https://doi.org/10.32871/rmrj2412.02.07

Longpre, S., Perisetla, K. Chen, A., Ramesh, N., DuBois, C., Singh, S. (2021). Entity-based knowledge conflicts in question answering. ArXiv. https://doi.org/10.48550/arXiv.2109.05052

Malik, M., & Isik, L. (2023). Relational visual representations underlie human social interaction recognition. Nature Communications, 14(7317), 1–11. https://doi.org/10.1038/s41467-023-43156-8

Mignot, E. (2012). The conceptualization of natural gender in English. Anglophonia: French Journal of English Studies, 16(32), 39-61. https://doi.org/10.4000/anglophonia.140

Mojapelo, M. (2023). Aspects of referent tracking in Northern Sotho. Sa B. Achiri-Taboh (Pat.), The Bantu Noun Phrase: Issues and Perspectives (pp. 210-226). London: Routledge. http://dx.doi.org/10.4324/9781003254188-13

Montalan, J., Layacan, J., Africa, D., et al. (2025). Batayan: A Filipino NLP benchmark for evaluating large language models arXiv. https://arxiv.org/abs/2502.14911

Nagaya, N. (2007). Information structure and constituent order in Tagalog. Languge and Linguistics, 8(1), 343-372.

Nordquist, R. (2019, Hulyo 3). Definitions and examples of hypernyms in English. ThoughtCo. https://www.thoughtco.com/hypernym-words-term-1690943

Nordquist, R. (2024, Abril 30). What are hyponyms in English? ThoughtCo. https://www.thoughtco.com/hyponym-words-term-1690946

Nordquist, R. (2025, Mayo 11). Definitions and examples of postmodifiers in English grammar. ThoughtCo. https://www.thoughtco.com/postmodifier-grammar-1691519

Oppenlaender, J. (2022). The creativity of text-to-image generation. In Proceedings of the 25th International Academic Mindtrek Conference (Academic Mindtrek ’22) (p. 11). ACM. https://doi.org/10.1145/3569219.3569352

Oppenlaender, J., Linder, R., & Silvennoinen, J. (2024). Prompting AI art: An investigation into the creative skill of prompt engineering. International Journal of Human–Computer Interaction, 1–23. https://doi.org/10.1080/10447318.2024.2431761

Reid, L., & Liao, H. (2004). A brief syntactic typology of Philippine languages. Language and Linguistics, 5(2), 433-490. http://hdl.handle.net/10125/32991

Souza, M. E., & Weigang, L. (2025). Grok, Gemini, ChatGPT, and Deepseek: Comparative and applications in conversational artificial intelligence. Inteligencia Artificial, 2(1), 1–7. http://dx.doi.org/10.5281/zenodo.14885243

Sperber, D., & Wilson, D. (2004). Relevance theory. Sa L. Horn & G. Ward (mga Pat.), Blackwell’s Handbook of Pragmatics (pp. 607-632). Blackwell.

Trojovsky, P., & Naveen, P. (2024). Overview and challenges of machine translation for contextually appropriate translations. iScience, 27(10), 1–25. https://doi.org/10.1016/j.isci.2024.110878

Villemaire, S. (2025). Canva. https://www.canva.com/learn/how-to-convert-text-images-ai-magic/.

Wang, Y., & Zhang, G. (2025). Lightweight text-to-image generation model based on contrastive language-image pre-training embeddings and conditional variational autoencoders. Electronics, 14(11), 1–31. https://doi.org/10.3390/electronics14112185

Warr, M. (2024). Beat bias? Personalization, bias, and generative AI. Sa J. Cohen & G. Solano (Eds.), Proceedings of Society for Information Technology and Teacher Education International Conference (pp. 1481–1488).

Downloads

Published

2026-07-30

How to Cite

Bosque, A. . (2026). Pragmatic Analysis of AI Image-to-Text Generation: The Case of ChatGP. Social Sciences and Development Review, 18(1), 277-298. https://doi.org/10.70922/x2sdfx87