¿Autodiseño? Evaluaciones de Modelos de Lenguaje de Gran Escala para Asistentes de UX
Revisión y Directrices
DOI:
https://doi.org/10.29147/datjournal.v11i2.1034Palabras clave:
LLM, Evaluación, Diseño, Investigación en UX, Redacción en UXResumen
Los modelos de lenguaje de gran escala (LLMs) están siendo rápidamente integrados en los flujos de trabajo de Design Systems, investigación en UX y redacción en UX, pero faltan métodos sistemáticos para evaluar su seguridad y utilidad. Basado en una revisión de alcance, este artículo mapea los principales desafíos de evaluación y propone tres marcos específicos para estas tareas. Cada marco conecta actividades de diseño con criterios de éxito medibles (p. ej., ≤ 3% de alucinación, < 2s de latencia) y sugiere conjuntos de datos para pruebas reproducibles. Una demostración con GPT-4o y Gemini 2.5 ilustra cómo las métricas identifican las fortalezas y debilidades de los modelos. Las directrices ofrecen a los equipos de diseño una hoja de ruta práctica para incorporar pruebas continuas y basadas en evidencia, promoviendo mayor rigor y responsabilidad en la práctica del diseño asistido por IA.
Descargas
Citas
BENDER, Emily M. et al. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: ACM CONFERENCE ON FAIRNESS, ACCOUNTABILITY, AND TRANSPARENCY (FAccT), 2021. Virtual Event, Canada. Proceedings [...]. New York: ACM, 2021. p. 610–623. DOI: 10.1145/3442188.3445922.
BROWN, Tim. Design Thinking: uma metodologia poderosa para decretar o fim das velhas ideias. Rio de Janeiro: Alta Books, 2020.
COSKUN, M. N. K. Spotify’s latest layoff is an opportunity to reconsider how we look at UX research. Fast Company, 2023. Available at: https://www.fastcompany.com/90999717/spotify-latest-layoff-ux-research. Accessed on: Jun. 25, 2025.
CRUTH, M. Discover the Spotify model. Atlassian, 2022. Available at: https://www.atlassian.com/agile/agile-at-scale/spotify. Accessed on: Jun. 25, 2025.
DIETZ, Laura et al. LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2504.19076.
FESSENDEN, T. Design Systems 101. Nielsen Norman Group, 2021. Available at: https://www.nngroup.com/articles/design-systems-101/. Accessed on: Jun. 26, 2025.
GAO, Mingqi et al. LLM-based NLG Evaluation: Current Status and Challenges. 2024. Preprint. DOI: https://doi.org/10.48550/arXiv.2402.01383.
GOTHELF, Jeff; SEIDEN, Josh. Lean UX: projetando ótimos produtos com times ágeis. 3. ed. Rio de Janeiro: Novatec Editora, 2022.
GOYAL, Tanya; LI, Junyi Jessy; DURRETT, Greg. News Summarization and Evaluation. 2022. Preprint. DOI: https://doi.org/10.48550/arXiv.2209.12356.
HULLEY, Stephen B. et al. Delineando a pesquisa clínica. 4. ed. Porto Alegre: Artmed, 2015.
HUSAIN, Hamel. Your AI Product Needs Evals. Hamel.dev, 2024. Available at: https://hamel.dev/blog/posts/evals/#motivation. Accessed on: Jun. 11, 2025.
LIN, Stephanie; HILTON, Jacob; EVANS, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 60., 2022, Dublin. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2022. p. 3214–3252. DOI: https://doi.org/10.18653/v1/2022.acl-long.229.
LIANG, Percy et al. Holistic Evaluation of Language Models. 2022. Preprint. Available at: https://arxiv.org/abs/2211.09110. Accessed on: Jun. 25, 2025.
MIZRAHI, Moran et al. State of What Art? A Call for Multi-Prompt LLM Evaluation. Transactions of the Association for Computational Linguistics, Cambridge, MA, v. 12, p. 933–949, 2024. DOI: https://doi.org/10.1162/tacl_a_00681.
MORAN, Kate. Usability (User) Testing. Nielsen Norman Group, 2019. Available at: https://www.nngroup.com/articles/usability-testing-101/. Accessed on: Jul. 18, 2025.
______. Calculating ROI for Design Projects in 4 Steps. Nielsen Norman Group, 2020. Available at: https://www.nngroup.com/articles/calculating-roi-design-projects/. Accessed on: Aug. 26, 2025.
______. The Four Dimensions of Tone of Voice. Nielsen Norman Group, 2023. Available at: https://www.nngroup.com/articles/tone-of-voice-dimensions/. Accessed on: Jun. 26, 2025.
NG, Andrew. Machine Learning Engineering for Production (MLOps) Specialization: Course 1. Coursera; DeepLearning.AI, 2025. Online course. Available at: https://www.coursera.org/learn/introduction-to-machine-learning-in-production. Accessed on: Jun. 25, 2025.
NIELSEN, Jakob. Usability 101. Nielsen Norman Group, 2012. Available at: https://www.nngroup.com/articles/usability-101-introduction-to-usability/. Accessed on: Jul. 18, 2025.
NIJS, Diane. Lens: a big shift in science – seeing change and innovation as a matter of emergence. In: ______ (Ed.). Advanced Imagineering: Designing Innovation as Collective Creation. Cheltenham, UK: Edward Elgar Publishing, 2019. chap. 2, p. 27–55.
RIBEIRO, Marco Tulio et al. Beyond Accuracy: Behavioral Testing of NLP Models with Checklist. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 58., 2020. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2020. p. 4902–4912. DOI: https://doi.org/10.18653/v1/2020.acl-main.442.
RICHARDS, Julian; WESSEL, Merten. Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants. In: INTERNATIONAL WORKSHOP ON BOTS IN SOFTWARE ENGINEERING (BOTSE), 6., 2025, Ottawa. Proceedings [...]. IEEE/ACM, 2025. (In press). Available at: https://arxiv.org/abs/2502.07956. Accessed on: Aug. 28, 2025.
SAURO, Jeff; LEWIS, James R. Quantifying the User Experience: Practical Statistics for User Research. 2. ed. Cambridge, MA: Morgan Kaufmann, 2016.
SCHULER, Douglas; NAMIOKA, Aki (Org.). Participatory Design: Principles and Practices. Hillsdale, NJ: Lawrence Erlbaum Associates, 1993.
SCHWARTZ, Reva et al. Reality Check: A New Evaluation Ecosystem is Necessary to Understand AI's Real World Effects. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2505.18893.
SPONHEIM, Christina. What Is User Research? Nielsen Norman Group, 2024. 1 video (2 min 37 s). Available at: https://www.nngroup.com/videos/what-is-user-research/. Accessed on: Jun. 26, 2025.
TRICCO, Andrea C. et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, Philadelphia, PA, v. 169, n. 7, p. 467–473, Oct. 2018. DOI: https://doi.org/10.7326/M18-0850.
WANG, Xuezhi et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. Preprint. Available at: https://doi.org/10.48550/arXiv.2203.11171. Accessed on: Aug. 28, 2025.
WOOD, Ian. Metadesign: Designing in the Anthropocene. London: Routledge, 2020.
ZHANG, Ziyin et al. Rethinking the A/B Test for Large Language Model-based Features: A Case Study from GitHub Copilot. In: ACM JOINT EUROPEAN SOFTWARE ENGINEERING CONFERENCE AND SYMPOSIUM ON THE FOUNDATIONS OF SOFTWARE ENGINEERING (ESEC/FSE), 31., 2023, San Francisco. Proceedings [...]. New York: ACM, 2023. p. 1000–1012. DOI: https://doi.org/10.1145/3611643.3616319.
Descargas
Publicado
Cómo citar
Número
Sección
Licencia
Derechos de autor 2026 DAT Journal

Esta obra está bajo una licencia internacional Creative Commons Atribución 4.0.






















