This is an outdated version published on 2025-06-29. Read the most recent version.

Visual ideation mediated by generative artificial intelligence: Lexical diversity, prompt complexity, and generation settings in DiffusionDB

Authors

  • MARIA FERNANDA RODRIGUEZ-RUIZ UNIVERSIDAD DE GUAYAQUIL Author

DOI:

https://doi.org/10.64747/2mgg3g14

Keywords:

generative artificial intelligence, visual ideation, prompt engineering, computational creativity, content analysis, DiffusionDB

Abstract

Prompt writing is an observable stage of ideation in text-to-image generative systems; however, its diversity, recurrence, and relationship with generation settings must be examined without equating textual complexity with psychological creativity. This study characterized visual ideation expressed in public DiffusionDB prompts through lexical diversity, structural elaboration, recurrence, and associations with generation parameters. We conducted a cross-sectional observational study using secondary public data. All 2,000,000 DiffusionDB 2M records were audited. Text analyses included 1,999,397 non-empty prompts, 1,522,692 unique normalized formulations, and a deterministic sample of 100,000 prompts for lexical metrics. Prompts had a median of 21 words and five segments. Repeated formulations accounted for 23.84% of records. The most frequent explicit markers were medium or technique (57.35%), style or artist (55.28%), and quality or detail (45.46%). The elaboration index had a median of two categories, and 60.65% of records combined at least two visual resources. Correlations between prompt length and generation settings were small within users (r = 0.12–0.16) and moderate across users (rho = 0.33–0.51). Prompts operated as comparatively rich compositional assemblages while relying on recurrent formulas. These findings describe early Stable Diffusion ideation practices; they do not measure human creativity or visual quality.

References

Covington, M. A., & McFall, J. D. (2010). Cutting the Gordian knot: The moving-average type-token ratio (MATTR). Journal of Quantitative Linguistics, 17(2), 94–100. https://doi.org/10.1080/09296171003643098

De Rosa Palmini, M.-T., & Cetinic, E. (2024). Patterns of creativity: How user input shapes AI-generated visual diversity [Preprint]. arXiv. https://arxiv.org/abs/2410.06768

Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. https://doi.org/10.1126/sciadv.adn5290

Epstein, Z., Hertzmann, A., Akten, M., Farid, H., Fjeld, J., Frank, M. R., Groh, M., Herman, L., Leach, N., Mahari, R., Pentland, A. S., Russakovsky, O., Schroeder, H., & Smith, A. (2023). Art and the science of generative AI. Science, 380(6650), 1110–1111. https://doi.org/10.1126/science.adh4451

Feng, Y., Wang, X., Wong, K. K., Wang, S., Lu, Y., Zhu, M., Wang, B., & Chen, W. (2023). PromptMagician: Interactive prompt engineering for text-to-image creation. IEEE Transactions on Visualization and Computer Graphics, 1–11. https://doi.org/10.1109/TVCG.2023.3327168

Liu, V., & Chilton, L. B. (2022). Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (pp. 1–23). Association for Computing Machinery. https://doi.org/10.1145/3491102.3501825

McCarthy, P. M., & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42(2), 381–392. https://doi.org/10.3758/BRM.42.2.381

Oppenlaender, J., Linder, R., & Silvennoinen, J. (2024). Prompting AI art: An investigation into the creative skill of prompt engineering. International Journal of Human–Computer Interaction. Advance online publication. https://doi.org/10.1080/10447318.2024.2431761

Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents [Preprint]. arXiv. https://arxiv.org/abs/2204.06125

Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10674–10685). IEEE. https://doi.org/10.1109/CVPR52688.2022.01042

Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Seyed Ghasemipour, S. K., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., & Norouzi, M. (2022). Photorealistic text-to-image diffusion models with deep language understanding [Preprint]. arXiv. https://arxiv.org/abs/2205.11487

Wadinambiarachchi, S., Kelly, R. M., Pareek, S., Zhou, Q., & Velloso, E. (2024). The effects of generative AI on design fixation and divergent thinking. In Proceedings of the CHI Conference on Human Factors in Computing Systems (pp. 1–18). Association for Computing Machinery. https://doi.org/10.1145/3613904.3642919

Wang, Z. J., Montoya, E., Munechika, D., Yang, H., Hoover, B., & Chau, D. H. (2022). DiffusionDB: A large-scale prompt gallery dataset for text-to-image generative models [Preprint]. arXiv. https://arxiv.org/abs/2210.14896

Zhou, E., & Lee, D. (2024). Generative artificial intelligence, human creativity, and art. PNAS Nexus, 3(3), pgae052. https://doi.org/10.1093/pnasnexus/pgae052

Downloads

Published

2025-06-29

Versions

Issue

Section

Research articles

How to Cite

Visual ideation mediated by generative artificial intelligence: Lexical diversity, prompt complexity, and generation settings in DiffusionDB. (2025). Sapiens Global, 1(1), 1-9. https://doi.org/10.64747/2mgg3g14