Current Bachelor Thesis Topics
Bachelor Topics WS 2026/2027
1. A Critical Review of the Ontology Use in Humanities & Social Science
Supervisors: Julia Schuller, Marta SabouKeywords: scoping literature review, critical interpretive synthesis, ontology adoption, critical data studies
Context:
Ontologies are used to agree on common vocabulary and relations of a certain domain. They enable data integration and interoperability across systems, management and agreement on shared knowledge, and obtaining and consistency checking of new knowledge by logical reasoning. In inherently structured contexts like proteins in biology or objects in factories, a construction of an ontology is a conducive abstraction and unification of a complex situation at hand. However, there is an increase in ontologies about rich concepts like trust [1], online harm [2], emotions [3] or human values [4]. Also, explicit calls form the humanities, and social sciences are expressed to increase ontology modelling efforts for computational analysis [5], [6], [7]. At the same time, the research field of science and technology studies (STS) has a long tradition flagging issues of the datafication of rich, qualitative phenomena like losing important contextual details or the solidification of stereotypes [8], [9], [10]. This research strand focusses mostly on tabular data while a critical perspective on ontologies remains largely unexplored [11].
Problem:
The use of computational methods in traditionally analogue disciplines has proven valuable, what social data science and digital humanities have shown. On a similar note, the usage and construction of ontologies for such research strands merits further exploration. However, to responsibly leverage that potential, guidelines and insights on best practices and known risks are missing.
Goal/expected results of the thesis:
This thesis develops a scoping review [12], [13] and understanding of which, how, and why ontologies are constructed and used in the humanities and social sciences. These descriptive findings are complemented with a critical synthesis [14] informed by existing literature in STS.
Research Questions:
How is the construction of ontologies in humanities and social sciences motivated?
Which challenges and limitations are described when constructing ontologies in humanities and social sciences?
How can underlying assumptions in the ontology construction be challenged from a critical data studies point of view?
Methodology & Approach:
Familiarise yourself with the usage of ontologies in humanities and social sciences in general
Familiarise yourself with criticism of digital classification and structuring
Systematically assemble a list of ontologies and related publications that match the focus of your choice
Extract and interpret information from the material you gathered
Prerequisites:
Prior experience in working with literature
Very good English skills
Interest in reading about, interpreting, and exploring unfamiliar topics (no worries, I am here to help, a comfortableness with larger amounts of text suffices)
References:
[1] G. Amaral, T. P. Sales, R. Baratella, D. Porello, R. Guizzardi, and G. Guizzardi, ‘ONTrust: A Reference Ontology of Trust’, Feb. 07, 2026, arXiv: arXiv:2602.07662. doi: 10.48550/arXiv.2602.07662.
[2] M. Banko, B. MacKeen, and L. Ray, ‘A Unified Taxonomy of Harmful Content’, in Proceedings of the Fourth Workshop on Online Abuse and Harms, S. Akiwowo, B. Vidgen, V. Prabhakaran, and Z. Waseem, Eds, Online: Association for Computational Linguistics, Nov. 2020, pp. 125–137. doi: 10.18653/v1/2020.alw-1.16.
[3] J. Hastings, W. Ceusters, B. Smith, and K. Mulligan, ‘The Emotion Ontology: Enabling Interdisciplinary Research in the Affective Sciences’, in Modeling and Using Context, Springer, Berlin, Heidelberg, 2011, pp. 119–123. doi: 10.1007/978-3-642-24279-3_14.
[4] S. De Giorgis, A. Gangemi, and R. Damiano, ‘Basic Human Values and Moral Foundations Theory in ValueNet Ontology’, in Knowledge Engineering and Knowledge Management, O. Corcho, L. Hollink, O. Kutz, N. Troquard, and F. J. Ekaputra, Eds, Cham: Springer International Publishing, 2022, pp. 3–18. doi: 10.1007/978-3-031-17105-5_1.
[5] T. Wolf et al., ‘Semantic Data for Humanities and Social Sciences (SDHSS): an Ecosystem of CIDOC CRM Extensions for Research Data Production and Reuse’, 2024, doi: 10.33968/9783966270502-05.
[6] T. Bosch and B. Zapilko, ‘Semantic Web Applications for the Social Sciences’, IASSIST Q., vol. 38, no. 4, p. 7, Dec. 2015, doi: 10.29173/iq896.
[7] M. Peponakis, S. Kapidakis, M. Doerr, and E. Tountasaki, ‘From Calculations to Reasoning: History, Trends and the Potential of Computational Ethnography and Computational Social Anthropology’, Soc. Sci. Comput. Rev., vol. 42, no. 1, pp. 84–102, Feb. 2024, doi: 10.1177/08944393231167692.
[8] G. C. Bowker and S. L. Star, Sorting Things Out: Classification and Its Consequences. The MIT Press, 1999. doi: 10.7551/mitpress/6352.001.0001.
[9] G. Neff and D. Nafus, Self-Tracking. MIT Press, 2016.
[10] S. Mau, The Metric Society: On the Quantification of the Social. John Wiley & Sons, 2019.
[11] D. Allhutter, ‘Of “Working Ontologists” and “High-Quality Human Components”: The Politics of Semantic Infrastructures’, in digitalSTS, Princeton University Press, 2019, pp. 326–348. doi: 10.1515/9780691190600-023.
[12] M. T. Pham, A. Rajić, J. D. Greig, J. M. Sargeant, A. Papadopoulos, and S. A. McEwen, ‘A scoping review of scoping reviews: advancing the approach and enhancing the consistency’, Res. Synth. Methods, vol. 5, no. 4, pp. 371–385, 2014, doi: 10.1002/jrsm.1123.
[13] D. Pollock et al., ‘“How-to”: scoping review?’, J. Clin. Epidemiol., vol. 176, p. 111572, Dec. 2024, doi: 10.1016/j.jclinepi.2024.111572.
[14] M. Dixon-Woods et al., ‘Conducting a critical interpretive synthesis of the literature on access to healthcare by vulnerable groups’, BMC Med. Res. Methodol., vol. 6, no. 1, p. 35, Jul. 2006, doi: 10.1186/1471-2288-6-35.
2. A Taxonomy of Self-Doubt Caused By Digital Technology
Supervisors: Julia Schuller, Marta SabouKeywords: taxonomy development, self-doubt, insecurity, uncertainty, literacy
Context:
The active usage of and passive consumptions or exposure to digital technology can affect individuals negatively in how they perceive themselves. One category of such negative self-perception is self-doubt provisionally defined as a feeling of uncertainty about one’s competence or capacity to make judgements [1], [2], [3]. In literature, this manifests in a variety of concepts. In the context of having the skills to use technology, research on literacy [4], [5] and the lack thereof especially of vulnerable groups like elderly [6] or disabled people [7] highlights the need for accessibility and self-efficacy consideration in technology design [7], [8]. Reversely, effects on the sense of self in the professional context caused by concerns regarding skill-erosion, future of professional role or expectations with regards to using tools when technology starts mastering one’s originally human expertise is studied [9], [10], [11]. In terms of the content and information technology produces and enables to consume, insecurities in discerning AI generated or from human generated ones [12] or the impact of the content on one’s self-image and resulting demands for self-presentation [13], [14] becomes relevant.
Problem:
Self-doubt as concept on its own, however, is not broadly studied in the context of digital technology besides few exceptions (compare [15], [16], [17], [18], [19]). Several related, more precise variants and facets are, however, prevalent. They, however, lack a shared theoretical grounding that would tie them together as one research area allowing for exchange of insights and ideas.
Goal/expected results of the thesis:
This thesis develops a taxonomy [20] for the concept of self-doubt caused by digital technology. The identified facets and unified vocabulary span and thereby unify dispersed strands in literature which allows for mutual enrichment and a common basis on which future research can build.
Research Questions:
Which kinds of self-doubts are induced by digital technology?
Which causes of self-doubts originate in digital technology?
Which consequences does self-doubt caused by digital technology entail?
Methodology & Approach:
Understand existing facets and dimensions of self-doubt in literature
Decide which aspects of self-doubt and digital technologies you want to focus on
Follow a structured approach for taxonomy building
(For master’s thesis: Analyse a concrete application along the taxonomy you developed and derive design guidelines to reduce the causes of self-doubt.)
Prerequisites:
Previous experience in working with literature
Comfortableness with larger amounts of text
Interest in interdisciplinary work
References:
[1] M. D. Braslow, J. Guerrettaz, R. M. Arkin, and K. C. Oleson, ‘Self-Doubt’, Soc. Personal. Psychol. Compass, vol. 6, no. 6, pp. 470–482, Jun. 2012, doi: 10.1111/j.1751-9004.2012.00441.x.
[2] A. D. Hermann, G. J. Leonardelli, and R. M. Arkin, ‘Self-Doubt and Self-Esteem: A Threat from within’, Pers. Soc. Psychol. Bull., vol. 28, no. 3, pp. 395–408, Mar. 2002, doi: 10.1177/0146167202286010.
[3] H. L. Mirels, P. Greblo, and J. B. Dean, ‘Judgmental self-doubt: beliefs about one’s judgmental prowess’, Personal. Individ. Differ., vol. 33, no. 5, pp. 741–758, Oct. 2002, doi: 10.1016/S0191-8869(01)00189-1.
[4] D. T. K. Ng, J. K. L. Leung, S. K. W. Chu, and M. S. Qiao, ‘Conceptualizing AI literacy: An exploratory review’, Comput. Educ. Artif. Intell., vol. 2, p. 100041, 2021, doi: 10.1016/j.caeai.2021.100041.
[5] D. Long and B. Magerko, ‘What is AI Literacy? Competencies and Design Considerations’, in Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, in CHI ’20. New York, NY, USA: Association for Computing Machinery, Apr. 2020, pp. 1–16. doi: 10.1145/3313831.3376727.
[6] H.-N. Kim, P. P. Freddolino, and C. Greenhow, ‘Older adults’ technology anxiety as a barrier to digital inclusion: a scoping review’, Educ. Gerontol., vol. 49, no. 12, pp. 1021–1038, Dec. 2023, doi: 10.1080/03601277.2023.2202080.
[7] E.-Y. Park, ‘Factors related to digital literacy in people with disabilities: Focus on self-efficacy and attitude toward digital devices and technology’, Int. Rev. Econ. Finance, vol. 104, p. 104671, Dec. 2025, doi: 10.1016/j.iref.2025.104671.
[8] J. B. Thatcher and P. L. Perrewé, ‘An Empirical Examination of Individual Traits as Antecedents to Computer Anxiety and Computer Self-Efficacy’, Manag. Inf. Syst. Q., vol. 26, no. 4, pp. 381–396, Dec. 2002, doi: 10.2307/4132314.
[9] M. K. K. Rony, Mst. R. Parvin, Md. Wahiduzzaman, M. Debnath, S. D. Bala, and I. Kayesh, ‘“I Wonder if my Years of Training and Expertise Will be Devalued by Machines”: Concerns About the Replacement of Medical Professionals by Artificial Intelligence’, Sage Open Nurs., vol. 10, p. 23779608241245220, Apr. 2024, doi: 10.1177/23779608241245220.
[10] J. Faulconbridge, A. Sarwar, and M. Spring, ‘How Professionals Adapt to Artificial Intelligence: The Role of Intertwined Boundary Work’, J. Manag. Stud., vol. 62, no. 5, pp. 1991–2024, 2025, doi: 10.1111/joms.12936.
[11] A. Kulal, ‘Why researchers hesitate to disclose AI use: development and validation of a multidimensional scale’, Ethics Behav., pp. 1–18, Apr. 2026, doi: 10.1080/10508422.2026.2661699.
[12] S. Choi, ‘The Modality-Congruent Carryover Effect: How Difficulty in Identifying (Deep)Fake News Impacts Self-Confidence in Truth Discernment and Susceptibility to Subsequent Disinformation’, Journal. Mass Commun. Q., p. 10776990251413726, Feb. 2026, doi: 10.1177/10776990251413726.
[13] J. R. Rui and M. A. Stefanone, ‘Strategic Image Management Online: Self-presentation, self-esteem and social network perspectives’, Inf. Commun. Soc., vol. 16, no. 8, pp. 1286–1305, Oct. 2013, doi: 10.1080/1369118X.2013.763834.
[14] T. H. H. Chua and L. Chang, ‘Follow me and like my beautiful selfies: Singapore teenage girls’ engagement in self-presentation and peer comparison on social media’, Comput. Hum. Behav., vol. 55, pp. 190–197, Feb. 2016, doi: 10.1016/j.chb.2015.09.011.
[15] A. Jamshaid and M. Grantham, ‘From imposter to algorithmic attribution: A theoretical model of AI-induced professional self-doubt’, Comput. Hum. Behav. Artif. Hum., vol. 8, p. 100326, May 2026, doi: 10.1016/j.chbah.2026.100326.
[16] P. Sanjeeva Kumar, ‘Technostress: A comprehensive literature review on dimensions, impacts, and management strategies’, Comput. Hum. Behav. Rep., vol. 16, p. 100475, Dec. 2024, doi: 10.1016/j.chbr.2024.100475.
[17] I. Nastjuk, S. Trang, J.-V. Grummeck-Braamt, M. T. P. Adam, and M. Tarafdar, ‘Integrating and synthesising technostress research: a meta-analysis on technostress creators, outcomes, and IS usage contexts’, Eur. J. Inf. Syst., vol. 33, no. 3, pp. 361–382, May 2024, doi: 10.1080/0960085X.2022.2154712.
[18] M. Y. Kayar, A. B. Körün, Ü. U. Topsakal, and S. A. Satıcı, ‘From Fear of Innovation to AI Dependency: the Mediating Roles of Fear of Failure and Self-Doubt’, Int. J. Ment. Health Addict., Nov. 2025, doi: 10.1007/s11469-025-01598-9.
[19] A. Asisof, ‘Retrieving Under Uncertainty: Towards a Chatbot Uncertainty Taxonomy (CUT) for Information Retrieval’, in Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR), in ICTIR ’25. New York, NY, USA: Association for Computing Machinery, Jul. 2025, pp. 155–166. doi: 10.1145/3731120.3744580.
[20] R. C. Nickerson, U. Varshney, and J. Muntermann, ‘A method for taxonomy development and its application in information systems’, Eur. J. Inf. Syst., vol. 22, no. 3, pp. 336–359, May 2013, doi: 10.1057/ejis.2012.26.
3. A Comparison of Responsible Requirements Engineering and Design Approaches
Supervisors: Julia Schuller, Marta SabouKeywords: value-sensitive design, value-based engineering, ethical engineering, scientific traditions
Context:
Technology has a profound impact on individuals, society, and other living species and the environment. Ideally that impact is responsibility steered to not only prevent harm but to additionally achieve a benefit [1], [2]. Manifold approaches ranging from abstract principles to concrete algorithms aim to facilitate responsible development for engineers and other project stakeholders. These cover the majority of steps along common software development lifecycles. At the beginning of the development process not only questions of functionality but also of non-functional requirements and considerations arise. Several major research traditions address that step of the development lifecycle either exclusively or as part of a more holistic responsible engineering process. Design science research proposes different traditions. Participatory [3], [4], [5] and critical design [6], [7] assume the idea that inclusion of various humans addresses their intentions best and that power dynamics require explicit consideration. Value-sensitive design focuses on the elicitation and realisation of human values explicitly. AI ethics researchers have so far mostly identified five central ethical principles that are now translated into technical reality, though not without caveats [1], [8], [9], [10]. Ethical engineering and design approaches on the abstract level of the entire design process have focussed on iterative and reflective practises [11], [12], [13], [14] and value-based engineering that enables standardised and operationalised development of ethics by design .
Problem:
Different engineering and design traditions suggest various paths that lead from abstract ideas and goals of responsible technology to the concrete realisation of such during development. Software developers and computer science researchers are increasingly expected by voices in the scientific and public debate but also by funding agencies and publishers to follow such ethical calls. However, industry mostly adheres to legal standards [8], [15] only while developers in general report facing challenges and barriers when aiming to realise responsible engineering in practice [16].
Goal/expected results of the thesis:
One step towards paving the way for those developers who are either genuinely interested in responsible engineering or enabling those who must address ethical expectations to not resort to a simple box-ticking exercise [17], [18] is to provide a clear, comparative overview of different approaches. This thesis provides a map and comparison of the major traditions of responsible design and requirements engineering so that technical experts can find a quick entry point to the field and informedly chose an approach that suits their project.
Research Questions:
Along which criteria and how do the main traditions of responsible design and requirements engineering differ?
Which assumptions and goals are embedded in the main traditions of responsible design and requirements engineering?
How can developers choose a responsible design or requirements engineering approach that is suitable for their project?
Methodology & Approach:
Familiarise yourself with the general literature on responsible and ethical technology
Choose the main traditions you want to look at and one to three major publications in that area
Qualitatively analyse [19] the chosen publications and derive high-level categories that describe the differences and assumptions
Highlight advantages and limitations, and develop specific questions developers can answer when drawing from that tradition
Prerequisites:
Familiarity with software engineering lifecycle and especially the requirements engineering stage
Interest in interdisciplinary work
Openness to critical reflection and ambiguous results
References:
[1] B. D. Mittelstadt, P. Allo, M. Taddeo, S. Wachter, and L. Floridi, ‘The ethics of algorithms: Mapping the debate’, Big Data Soc., vol. 3, no. 2, p. 2053951716679679, Dec. 2016, doi: 10.1177/2053951716679679.
[2] S. Spiekermann et al., ‘Values and Ethics in Information Systems: A State-of-the-Art Analysis and Avenues for Future Research’, Bus. Inf. Syst. Eng., vol. 64, no. 2, pp. 247–264, Apr. 2022, doi: 10.1007/s12599-021-00734-8.
[3] D. Schuler and A. Namioka, Participatory Design: Principles and Practices. CRC Press, 1993.
[4] J. Simonsen and T. Robertson, Routledge International Handbook of Participatory Design. Routledge, 2013.
[5] K. Halskov and N. B. Hansen, ‘The diversity of participatory design research practice at PDC 2002–2012’, Int. J. Hum.-Comput. Stud., vol. 74, pp. 81–92, Feb. 2015, doi: 10.1016/j.ijhcs.2014.09.003.
[6] J. Bardzell and S. Bardzell, ‘What is “critical” about critical design?’, in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, in CHI ’13. New York, NY, USA: Association for Computing Machinery, Apr. 2013, pp. 3297–3306. doi: 10.1145/2470654.2466451.
[7] S. Costanza-Chock, Design Justice: Community-Led Practices to Build the Worlds We Need. Erscheinungsort nicht ermittelbar: The MIT Press, 2020.
[8] T. Hagendorff, ‘The Ethics of AI Ethics: An Evaluation of Guidelines’, Minds Mach., vol. 30, no. 1, pp. 99–120, Mar. 2020, doi: 10.1007/s11023-020-09517-8.
[9] J. Morley, L. Floridi, L. Kinsey, and A. Elhalal, ‘From What to How: An Initial Review of Publicly Available AI Ethics Tools, Methods and Research to Translate Principles into Practices’, Sci. Eng. Ethics, vol. 26, no. 4, pp. 2141–2168, Aug. 2020, doi: 10.1007/s11948-019-00165-5.
[10] J. Morley, L. Kinsey, A. Elhalal, F. Garcia, M. Ziosi, and L. Floridi, ‘Operationalising AI ethics: barriers, enablers and next steps’, AI Soc., vol. 38, no. 1, pp. 411–423, Feb. 2021, doi: 10.1007/s00146-021-01308-8.
[11] I. van de Poel, ‘Translating Values into Design Requirements’, in Philosophy of Engineering and Technology, vol. 15, Springer Nature, 2013, pp. 253–266. doi: 10.1007/978-94-007-7762-0_20.
[12] J. Gogoll, N. Zuber, S. Kacianka, T. Greger, A. Pretschner, and J. Nida-Rümelin, ‘Ethics in the Software Development Process: from Codes of Conduct to Ethical Deliberation’, Philos. Technol., vol. 34, no. 4, pp. 1085–1108, Dec. 2021, doi: 10.1007/s13347-021-00451-w.
[13] N. Zuber, J. Gogoll, S. Kacianka, A. Pretschner, and J. Nida-Rümelin, ‘Empowered and embedded: ethics and agile processes’, Humanit. Soc. Sci. Commun., vol. 9, no. 1, p. 191, Jun. 2022, doi: 10.1057/s41599-022-01206-4.
[14] I. van de Poel and L. Royakkers, Ethics, Technology, and Engineering: An Introduction. Hoboken, USA: Wiley-Blackwell, 2023.
[15] V. Vakkuri, K.-K. Kemell, M. Jantunen, and P. Abrahamsson, ‘“This is just a prototype”: How ethics are ignored in software startup-like environments’, in International Conference on Agile Software Development, Springer International Publishing Cham, 2020, pp. 195–210. Accessed: Aug. 14, 2025. [Online].
[16] B. Rakova, J. Yang, H. Cramer, and R. Chowdhury, ‘Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for Shifting Organizational Practices’, Proc ACM Hum-Comput Interact, vol. 5, no. CSCW1, p. 7:1-7:23, Apr. 2021, doi: 10.1145/3449081.
[17] D. Greene, A. L. Hoffmann, and L. Stark, ‘Better, Nicer, Clearer, Fairer: A Critical Assessment of the Movement for Ethical Artificial Intelligence and Machine Learning’, Hawaii Int. Conf. Syst. Sci. 2019 HICSS-52, Jan. 2019, [Online]. Available: aisel.aisnet.org/hicss-52/dsm/critical_and_ethical_studies/2
[18] A. H. Kiran, N. Oudshoorn, and P.-P. Verbeek, ‘Beyond checklists: toward an ethical-constructive technology assessment’, J. Responsible Innov., vol. 2, no. 1, pp. 5–19, Jan. 2015, doi: 10.1080/23299460.2014.992769.
[19] A. Lacey and D. Luff, ‘Qualitative Data Analysis’, NIHR RDS East Midl..
4. Text Mining and Machine Learning
Supervisor: Johann MitlöhnerText Mining and Machine Learning
Text mining aims to turn written natural language into structured data that allow for various types of analysis which are hard or impossible on the text itself; machine learning aims to automate the process using a variety of adaptive methods, such as artificial neural nets which learn from training data. Typical goals of text mining are Classification, Sentiment Detection, and other types of Information Extraction, e.g. Named Entity Recognition: identify people, places, organizations; Relation Extraction, e.g. locations of organizations.
Connectionist methods and deep learning in particular have achieved much attention and success recently; these methods tend to work well on large training datasets which require ample computing power. Our institute can provide access to high performance GPU units for student use in thesis projects. It is recommended to use a framework such as PyTorch or Tensorflow/Keras for developing your deep learning application; the changes required to go from CPU to GPU computing will be minimal. This means that you can start developing using your PC, with a small subset of the training data; when you later transition to the GPU server more performance will mean that larger datasets become feasible.
Some References:
Minqing Hu, Bing Liu: Mining and summarizing customer reviews. KDD '04: Proceedings of the tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 168-177, ACM, 2004
For a more recent work and overview e.g.: Percha B. Modern Clinical Text Mining: A Guide and Review. Annu Rev Biomed Data Sci. 2021 Jul 20;4:165-187. doi: 10.1146/annurev-biodatasci-030421-030931. Epub 2021 May 26. PMID: 34465177.
Datasets can be found e.g. at huggingface and kaggle.
keywords: artificial neural networks, machine learning, text mining
5. Visualizing Data in Virtual and Augmented Reality
Supervisor: Johann MitlöhnerVisualizing Data in Virtual and Augmented Reality
How can AR and VR be used to improve exploration of data? Developing new methods for exploring and analyzing data in virtual and augmented reality presents many opportunities and challenges, both in terms of software development and design inspiration. There are various hardware options, starting with Google Cardboard, to more sophisticated and expensive, such as Quest, and many others. Taking part in this challenge demands programming skills as well as creativity. A basic VR or AR application for exploring a specific type of (open) data will be developed by the student. The use of a platform-independent kit such as A-Frame is essential, as the application will be compared in a small user study to its non-VR version in order to identify advantages and disadvantages of the visualization method implemented. Details will be discussed with supervisor.
Some References:
Butcher, Peter WS, and Panagiotis D. Ritsos. "Building Immersive Data Visualizations for the Web." Proceedings of International Conference on Cyberworlds (CW'17), Chester, UK. 2017.
Teo, Theophilus, et al. "Data fragment: Virtual reality for viewing and querying large image sets." Virtual Reality (VR), 2017 IEEE. IEEE, 2017.
Millais, Patrick, Simon L. Jones, and Ryan Kelly. "Exploring Data in Virtual Reality: Comparisons with 2D Data Visualizations." Extended Abstracts of the 2018 CHI Conference on Human Factors in Computing Systems. ACM, 2018.
Yu Shu, Yen-Zhang Huang, Shu-Hsuan Chang, and Mu-Yen Chen (2019). Do virtual reality head-mounted displays make a difference? a comparison of presence and self-efficacy between head-mounted displays and desktop computer-facilitated virtual environments. Virtual Reality, 23(4):437-446.
Korkut, E. H., and Surer, E. (2023). Visualization in virtual reality: a systematic review. Virtual Reality, 27(2), 1447-1480.
keywords: virtual reality, augmented reality, data visualization, data exploration
6. Embracing Diverse Ways of Knowing: Exploring Neurodiversity and Knowledge Management in Organizations
Supervisors: Florian Kragulj, Sophia MayrhoferNeurodivergent employees bring distinctive perspectives, experiences, and ways of knowing to organizations. Yet organizational norms often privilege neurotypical forms of communication, interaction, and contribution, leaving valuable knowledge unrecognized or excluded from organizational processes. A neurodiversity perspective challenges these norms by understanding neurological differences – including autism, ADHD, dyslexia, DCD, and Tourette syndrome – as part of human diversity rather than as deficiencies. This raises important questions for knowledge management about how different ways of thinking and communicating influence managing knowledge within organizations.
This thesis invites you to bring together two research fields that have rarely been considered in combination: organizational neurodiversity and knowledge management. Through a systematic literature review, you will investigate how neurodiversity is defined and explained in organizational research. You will then explore the knowledge management literature to identify key knowledge-related processes and outcomes that are relevant to organizational neurodiversity.
Drawing on your findings, you are invited to develop a structured synthesis/preliminary conceptual framework showing how organizational approaches to neurodiversity might shape your chosen knowledge processes and outcomes. You are encouraged to think creatively, uncover relationships between the two fields, identify gaps in the existing literature, and suggest promising directions for future research.
For further questions, please contact Sophia Mayrhofer (sophia.mayrhofer@s.wu.ac.at).
Keywords: Neurodiversity in Organizations, Neuroinclusive Organizations, Neurodivergent Employees, Knowledge Management, Systematic Literature Review
Initial References:
Nonaka, I., Toyama, R., & Konno, N. (2000). SECI, Ba and leadership: A unified model of dynamic knowledge creation. Long Range Planning, 33(1), 5–34.
Doyle, N. (2020). Neurodiversity at work: A biopsychosocial model and the impact on working adults. British Medical Bulletin, 135(1), 108–125.
LeFevre-Levy, R., Melson-Silimon, A., Harmata, R., Hulett, A. L., & Carter, N. T. (2023). Neurodiversity in the workplace: Considering neuroatypicality as a form of diversity. Industrial and Organizational Psychology, 16(1), 1–19.
7. Knowledge Management Processes in Startups: What Do We Know Across Development Stages?
Supervisors: Florian Kragulj, Tiffany Ar RahimKnowledge is both the lifeblood and the Achilles' heel of startups. Startups operate under conditions of extreme uncertainty, severe resource constraints, and rapidly evolving market demands, forces that compel entrepreneurs to creatively leverage whatever knowledge and resources they have at hand. This phenomenon is known as entrepreneurial bricolage. Startups must continuously acquire, create, share, and apply knowledge to solve immediate problems and seize emerging opportunities.
Yet startups are far from static entities. As they develop and grow, their organizational landscape transforms dramatically, their resource availability expands, structures formalize, team sizes increase, and coordination mechanisms become more complex. A startup's knowledge management processes that work in a five-person team may become obsolete or ineffective when the team grows to fifty.
In this bachelor thesis, you will explore knowledge management processes in startups from a knowledge management perspective:
What knowledge management processes are identified in startups?
How are these knowledge management processes defined and distinguished in the literature?
How are these knowledge management processes described across different stages of startup development, and what patterns emerge?
The expected output of this study is a synthesis of findings into a structured framework that maps knowledge management processes in startups and shows how these processes evolve across development stages.
For further questions, please contact Tiffany Ar Rahim (sekretariat@dpkm.wu.ac.at).
Keywords: Knowledge Management Processes, Startups, Startup Development Stages, Structured Literature Review
Initial References:
Baker, T., & Nelson, R. E. (2005). Creating something from nothing: Resource construction through entrepreneurial bricolage. Administrative Science Quarterly, 50(3), 329–366.
Centobelli, P., Cerchione, R., & Esposito, E. (2017). Knowledge Management in Startups: Systematic Literature Review and Future Research Agenda. Sustainability, 9(3), 361. https://doi.org/10.3390/su9030361
Gaimon, C., & Bailey, J. (2013). Knowledge Management for the Entrepreneurial Venture. Production and Operations Management, 22(6), 1429-1438. https://doi.org/10.1111/j.1937-5956.2012.01337.x
Zobnina, M. (2015). Startup development, investments, and growth barriers. In Emerging Markets and the Future of the BRIC Nations (pp. 111–124). Edward Elgar Publishing.
8. Towards Detecting Malicious Cryptocurrency Transactions With Machine Learning
Supervisors: Jennifer-Marieclaire Sturlese, Marta SabouSupervising: Jennifer-Marieclaire Sturlese
Grading: Marta Sabou
Winter Semester 2026
Title: Towards Detecting Malicious Cryptocurrency Transactions With Machine Learning
Abstract:
Cryptocurrencies have created new opportunities for financial transactions while also enabling novel forms of cybercrime, including money laundering, ransomware, fraud, and further dark-web market activities. Blockchain mechanisms provide transparency through publicly accessible transaction records, yet the vast volume of transactions can make the identification of suspicious activities challenging. As a result, machine learning has emerged as promising tool for detecting malicious cryptocurrency transactions.
This bachelor thesis investigates blockchain security from a machine learning perspective. You will use the Elliptic Bitcoin Dataset, which contains labeled cryptocurrency transactions, and will explore patterns of suspicious transaction behavior. By employing supervised and unsupervised learning, you will aim towards evaluating the effectiveness of contemporary ML methods for identifying the most relevant features contributing to classification performance.
The thesis is structured into two parts: a literature review on blockchain technology and its security, cryptocurrencies and anti-money-laundering procedures related to cryptocurrencies (= theoretical part); an explanation of the research design using Python-based data analytics and machine learning tools; the application of this methodology involving data preprocessing, feature engineering, model development, visualization, and performance evaluation using the Elliptic Bitcoin Dataset (= empirical part). Your potential research question: To what extent can machine learning be used to identify malicious transactions based on cryptocurrency transaction data?
In order to receive this topic, you are required to demonstrate knowledge in: Data analysis with Python, machine learning packages (please state experience in motivation letter), basic understanding of blockchain concepts, completed K-5 of your specialization (Knowledge Management and Data Science, both welcomed, latter is preferred).
Keywords: blockchain security, cryptocurrency transactions, fraud detection
Initial References:
Taudes, A. (2024). Cryptoeconomic-Blockchains, Game Theory and Artificial Intelligence.
Wahrstätter, A., Taudes, A., & Svetinovic, D. (2024). Reducing privacy of coinjoin transactions: Quantitative bitcoin network analysis. IEEE Transactions on Dependable and Secure Computing, 21(5), 4543-4558.
Gomes, J., Khan, S., & Svetinovic, D. (2023). Fortifying the blockchain: A systematic review and classification of post-quantum consensus solutions for enhanced security and resilience. IEEE Access, 11, 74088-74100.
Lorenz, J., Neudert, M., Orlitzky, M., & Popp, T. (2020). Machine Learning Methods to Detect Money Laundering in the Bitcoin Blockchain in the Presence of Label Scarcity. Proceedings of the ICAIF Conference.
Elliptic Dataset https://www.kaggle.com/datasets/ellipticco/elliptic-data-set
9. Towards Interpretable Cybersecurity Threat Detection With Explainable Artificial Intelligence
Supervisors: Jennifer-Marieclaire Sturlese, Marta SabouSupervising: Jennifer-Marieclaire Sturlese
Grading: Marta Sabou
Winter Semester 2026
Title: Towards Interpretable Cybersecurity Threat Detection With Explainable Artificial Intelligence
Abstract:
Cyberattacks pose increasing challenges with regards to securing socio-technical systems. While machine learning has become an established approach for detecting malicious network activities, automated decisions from “black box” models are difficult to validate by the human-side. One reason for this is the lack of transparency that reduces trust in machine learning decisions and hinder their practical adoption in cybersecurity operations.
This bachelor thesis investigates cybersecurity threat detection from an explainable artificial intelligence (XAI) perspective. You will use an open dataset (either CICIDS2017, NSL-KDD, or UNSW-NB15) and explore patterns associated with normal and malicious network behavior. By employing supervised ML and explainability methods, you will aim towards evaluating the effectiveness of contemporary ML models and identify the most relevant features contributing to threat detection decisions.
The thesis is structured into two parts: a literature review on secure socio-technical systems, cybersecurity threat detection, and explainable artificial intelligence (= theoretical part); an explanation of the research design using Python ML and XAI tools; the application of this methodology involving data preprocessing, feature engineering, model development, explainability analysis, visualization, and performance evaluation using a publicly available network intrusion dataset (= empirical part). Your potential research question: To what extent can XAI support the interpretation of cybersecurity threat detection?
In order to receive this topic, you are required to demonstrate knowledge in: Data analysis with Python, machine learning packages (please state experience in motivation letter), basic understanding of cybersecurity concepts, completed K-5 of your specialization (Knowledge Management and Data Science, both welcomed, latter is preferred).
Keywords: cybersecurity, explainable AI, XAI, threat detection
Initial References:
Sharshar, M., Saber, A. M., Svetinovic, D., Youssef, A. M., Kundur, D., & El-Saadany, E. F. (2025, October). Large Language Model-Based Framework for Explainable Cyberattack Detection in Automatic Generation Control Systems. In 2025 IEEE Electrical Power and Energy Conference (EPEC) (pp. 424-429). IEEE.
Yaseen, A. (2023). AI-driven threat detection and response: A paradigm shift in cybersecurity. International Journal of Information and Cybersecurity, 7(12), 25-43.
Buczak, A. L., & Guven, E. (2016). A survey of data mining and machine learning methods for cybersecurity intrusion detection. IEEE Communications Surveys & Tutorials, 18(2), 1153–1176.
Datasets: https://www.kaggle.com/datasets/chethuhn/network-intrusion-dataset -or- https://www.kaggle.com/datasets/hassan06/nslkdd -or- https://www.kaggle.com/datasets/dhoogla/unswnb15
10. Facilitating Interdisciplinary Collaboration through Knowledge Graph-based Expertise Discovery and Recommendation: A Use Case of an Austrian University
Supervisors: Daniil Dobriy, Axel PolleresTitle: Facilitating Interdisciplinary Collaboration through Knowledge Graph-based Expertise Discovery and Recommendation: A Use Case of an Austrian University
Advisor: Daniil Dobriy
The thesis/topic can be written by more than one student
Other important information:
If your thesis involves software development, Git version control will be used. Prior Git experience is not mandatory - a training course will be provided if needed. Artifacts produced during your thesis should meet the following requirements whenever permissible:
| Documentation | Publication | License | |
| Software | StandardREADME[1] | Institute’s GitLab[2] | MIT license[3] |
| Datasets | DCAT[4] | WU data portal[5] | CC BY 4.0[6] |
| Other artifacts | Standard README | Institute’s GitLab | CC BY 4.0 |
For computationally intensive applications, students will be provided SSH access to institute’s server infrastructure. To facilitate deployment, complex applications are recommended to be containerized using Docker - training available if needed.
Thesis description:
Despite being transformative, integrative and accelerating innovation [3], interdisciplinary research is associated with a number or institutional, professional and organizational challenges [1]. Inspired by the theory of weak ties [4], this research topic aims at developing a comprehensive framework for a researcher discovery and expert recommendation [2] that leverages semantic web technologies [5] to enable interdisciplinary collaboration at academic institutions. Positioned at WU Vienna and using the research information management systems available at the university (e.g., Pure) as a practical case study, the thesis combines a variety of topics including the challenges of interdisciplinary collaboration as a literature and survey study with technical aspects such as ETL (WU research platforms), knowledge graph construction, ontology re-use and recommendation frameworks to solve a genuine institutional challenge.
This thesis topic envisions collaboration with the WU Research Service Center. The final scope and focus of the topic will be decided with the students based on the relevant expertise and research interests. The topic is multi-faceted and will accommodate multiple students - i.e., we encourage multiple students to apply and work in parallel on the evaluation of different aspects of the topic.
Prerequisites:
Familiarity with ETL (Extract, Transform, Load) pipelines and knowledge representation/knowledge graphs is advantageous.
Interest in developing thesis findings into an academic publication.
Strong time management skills and ability to meet project milestones. A structured workflow with clear milestones will guide you through the thesis development process.
References:
[1] Daniel, K. L., McConnell, M., Schuchardt, A., & Peffer, M. E. (2022). Challenges facing interdisciplinary researchers: Findings from a professional development workshop. Plos one, 17(4), e0267234.
[2] Nikzad–Khasmakhi, Narjes, M. A. Balafar, and M. Reza Feizi–Derakhshi. "The state-of-the-art in expert recommendation systems." Engineering Applications of Artificial Intelligence 82 (2019): 126-147.
[3] Fagan, Jesse, et al. "Assessing research collaboration through co-authorship network analysis." The journal of research administration 49.1 (2018): 76.
[4] Granovetter, Mark. "The strength of weak ties: A network theory revisited." Sociological theory (1983): 201-233.
[5] Hogan, Aidan, et al. "Knowledge graphs." ACM Computing Surveys (Csur) 54.4 (2021): 1-37.
Keywords: Expert Recommendation Systems, Interdisciplinary Research, Scientific Linked Data, Knowledge Graphs, Ontologies, Research Information Management Systems, Scientific Collaboration Networks
[1] See https://github.com/RichardLitt/standard-readme
[2] I.e., https://git.ai.wu.ac.at
[3] See https://opensource.org/license/mit
[4] See https://www.w3.org/TR/vocab-dcat-3/
[5] I.e., https://data.wu.ac.at/portal
[6] See https://creativecommons.org/licenses/by/4.0/deed.en
11. Developing and Evaluating Components of the Search Engine for the Web of Data (search.ai.wu.ac.at)
Supervisors: Daniil Dobriy, Axel PolleresTitle: Developing and Evaluating Components of the Search Engine for the Web of Data (search.ai.wu.ac.at)
Advisor: Daniil Dobriy
The thesis/topic can be written by more than one student
Other important information:
If your thesis involves software development, Git version control will be used. Prior Git experience is not mandatory - a training course will be provided if needed. Artifacts produced during your thesis must meet the following requirements whenever permissible:
| Documentation | Publication | License | |
| Software | StandardREADME[1] | Institute’s GitLab[2] | MIT license[3] |
| Datasets | DCAT[4] | WU’s data portal[5] | CC BY 4.0[6] |
| Other artifacts | Standard README | Institute’s GitLab | CC BY 4.0 |
For computationally intensive applications, students will be provided SSH access to institute’s server infrastructure. To facilitate deployment, complex applications are recommended to be containerized using Docker - training available if needed.
Thesis description:
Standard protocols such as the Model Context Protocol (MCP) [2] that allow LLMs to connect to tools, have recently boosted the applications of LLMs to develop "agentic" AI applications, which - powered by - LLMs' planning capabilities promise to solve complex tasks with the access of external tools and databases. Use cases explored so far in the literature have mostly focused on interactions with single (e.g. relational) databases, with tasks such as schema exploration, query formulation from natural language (text-to-SQL) and results translation to desired output formats being successfully delegated to such agents. On the contrary, SPARQL as a standard query language offers even more flexibility to combine various data sources through (a) endpoints readily implementing a standard protocol, (b) standardized metadata formats to self-describe an endpoints schema and capabilities potentially leveraging dynamic discovery, and (c) SPARQL's native capability to federate queries across multiple such endpoints.
Previous pipelines have been proposed for text-to-SPARQL generation and federated querying [3]. One of the proposed agentic SPARQL federation architectures (see Figure 1) relies on a centralized “Catalogue” that facilitates endpoint discover and schema exploration. Previously, SPARQLES [1] has been such proposed “endpoint monitoring service” and is a precursor to the Catalogue. In this thesis, your goal will be to explore types of metadata and strategies that could support (a) endpoint discovery (i.e. “searching for an applicable endpoint/knowledge graph that could deliver an answer to a specific natural language question or SPARQL query”) and (b) schema exploration (i.e. “retrieving the part of schema that is relevant to a particular natural language question or SPARQL query”). More detailed goal definition will be formulated with the student in the early stages of the thesis planning process.
The topic will accommodate multiple students – i.e., we encourage multiple students to apply and work in parallel on the evaluation of different strategies with the evaluation testbed we provide.
Prerequisites:
Familiarity with SPARQL and MCP is advantageous.
Interest in developing thesis findings into an academic publication.
Strong time management skills and ability to meet project milestones. A structured workflow with clear milestones will guide you through the thesis development process.
References:
Pierre-Yves Vandenbussche, Jürgen Umbrich, Luca Matteis, Aidan Hogan, and Carlos Buil-Aranda. 2017. SPARQLES: Monitoring public SPARQL endpoints. Semantic web 8, 6 (2017), 1049–1065.
Anthropic. 2025. Model Context Protocol Specification, Version 2025-06-18. MCP Specification. modelcontextprotocol.io/specification/2025-06-18 Accessed: 2025-09-14.
Emonet, V., Bolleman, J., Duvaud, S., de Farias, T. M., & Sima, A. C. (2024). Llm-based sparql query generation from natural language over federated knowledge graphs. arXiv preprint arXiv:2410.06062. https://arxiv.org/html/2410.06062v2
Keywords: Query Federation,Retrieval-Augmented Generation, Large Language Models, Named Entity Recognition, Knowledge Graphs, Knowledge Graph Question Answering (KGQA)
[1] See https://github.com/RichardLitt/standard-readme
[2] I.e., https://git.ai.wu.ac.at
[3] See https://opensource.org/license/mit
[4] See https://www.w3.org/TR/vocab-dcat-3/
[5] I.e., https://data.wu.ac.at/portal
[6] See https://creativecommons.org/licenses/by/4.0/deed.en
12. Pseudo-Relevance Feedback for N-Gram Generative Retrieval
Supervisors: Adrian Bracher, Svitlana VakulenkoBackground: Generative Retrieval is an emerging information retrieval technique by training an autoregressive language model to generate document identifiers (docids) directly. N-gram-based architectures, such as SEAL, use contiguous substrings extracted from corpus text as docids. However, empirical analysis shows that n-gram GR models frequently suffer from parametric biases, decoding errors and ambiguous, short docids. Pseudo-Relevance Feedback (PRF) is a classical retrieval technique that assumes top-ranked documents from an initial search contain relevant terms, which can be extracted and combined with the original query to improve subsequent retrieval performance. To overcome these limitations, this project explores a PRF framework for n-gram-indexed Generative Retrieval.
RQ: How can classical Pseudo-Relevance Feedback (PRF) techniques be adapted to n-gram-indexed generative retrieval?
References:
Bevilacqua et al., “Autoregressive search engines: Generating substrings as document identifiers.”, arxiv:2204.10628, 2022
Lavrenko and Croft., “Relevance-Based Language Models.”, doi.org/10.1145/383952.3839, 2001
Takacs et al., “Understanding and Debugging Failures in N-Gram-Based Generative Retrieval.”, arxiv:2606.17721, 2026
Keywords: Generative Retrieval, Pseudo-Relevance Feedback (PRF), Information Retrieval
Prior Knowledge:
Python programming skills (scripting, data structures).
Foundations of Information Retrieval & Data Science
Suitable for bachelor or master thesis
13. Failure Mode Analysis Generative Retrieval
Supervisors: Adrian Bracher, Svitlana VakulenkoBackground: Recent investigations into Generative Retrieval established a comprehensive taxonomy of failure modes across representation, training, inference, and response stages (https://arxiv.org/pdf/2606.17721). This work mapped out initial failure modes, but left important open questions regarding true causality and architectural generalizability. This thesis is about extending the analysis. For example, the work could be extended by controlling for confounding variables (e.g., query complexity) that may obscure the results. Another direction could be extending this failure evaluation framework to other methods, to examine alternative architectures (e.g., SPLADE), in order to better understand the failures and shortcomings across information retrieval paradigms.
RQs:
How do confounding variables, such as query complexity, impact the failure analysis?
How do identified GR failure modes generalize across alternative retrieval paradigms?
References:
Takacs et al., “Understanding and Debugging Failures in N-Gram-Based Generative Retrieval.”, arxiv:2606.17721, 2026
Bevilacqua et al., “Autoregressive search engines: Generating substrings as document identifiers.”, arxiv:2204.10628, 2022
Formal et al., “SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking”, arxiv:2107.05720, 2021
Keywords: Data Analysis, Information Retrieval, Failure Modes
Prior Knowledge:
Programming skills for data analysis (Python or R)
Suitable for bachelor thesis
14. Conversational Search for Cultural Archives
Supervisors: Adrian Bracher, Svitlana VakulenkoBackground: Digital cultural institutions (e.g., Wien Museum, Wienbibliothek, Vienna History Wiki) host large open-data, digital collections. However, traditional keyword-based catalog searches fail to support intuitive, exploratory discovery. In this project we build a chat assistant that allows users to search and explore Vienna's cultural archives interactively. We will collect archive data, and apply state-of-the-art conversational search methods, and rigorously evaluate how well the system works for this use-case.
RQ:
How can a conversational RAG architecture be applied to the domain of digital cultural collections?
References:
Jiang et al., “Orcheo: A Modular Full-Stack Platform for Conversational Search”, arxiv:2602.14710, 2026
Vakulenko et al., “Question Rewriting for Conversational Question Answering”, arxiv:2004.14652, 2020
Mo et al., “A Survey on Conversational Search”, arxiv:2410.15576, 2025
Keywords: Conversational Search, Retrieval-Augmented Generation (RAG), Cultural Archives, Multi-Turn Dialogue
Prior Knowledge:
Python programming skills
Familiarity with modern LLM frameworks (e.g., Hugging Face, Ollama) and basic NLP concepts
Suitable for bachelor or master thesis
15. Implementing an Agentic Workflow for Cybersecurity Defense
Supervisor: Elmar KieslingBackground:
Large Language Models are seeing widespread adoption in cybersecurity applications for both offensive [3] and defensive [5] purposes [8]. Whereas their potential for malicious use has raised widespread concern and media coverage, their potential to improve security is also widely recognized. This has sparked research interest into how generative AI systems can automate, support and/or scale tasks such as log analysis [1], vulnerability assessment and mitigation [4,7], threat detection [2], or attack graph construction [9]. More recently, cybersecurity has also emerged as a particularly promising application domain for agentic AI, where teams of autonomous AI agents are envisioned to solve complex tasks in distributed workflows. Although research is at a very early stage, initial results suggest that multi-agent workflows where agents use tools, perform multi-step reasoning, and ultimately support rapid decision making under pressure have strong potential to improve cybersecurity [6].
Research problem:
In this Bachelor thesis, you will as a preliminary step review promising opportunities for agentic AI to support analytic tasks in cybersecurity (e.g., vulnerability assessment, risk assessment, penetration testing, intrusion detection, threat hunting, impact assessment, root cause analysis, incident response etc.). Based on this initial survey, the thesis will focus on a selected task (or small subset of tasks) and a scenario for experimentation. It will document the implementation of an agentic workflow and evaluate and compare how agentic AI can support the selected analytic task.
Your thesis may address research questions such as:
What are key criteria when selecting an agentic framwork for cybersecurity analysis?
What are dimensions to compare available frameworks for cybersecurity analysis?
What are key challenges in the implementation of agentic workflows for cybersecurity analysis?
How does the implemented agentic workflow compare to a manual process?
What are specific benefits, risks, and challenges?
Required skills:
This topic is aimed at experimenting with agentic frameworks and implementing agentic workflow(s) in a selected cybersecurity scenario. This necessitates interest in developing the necessary skills to work with agentic frameworks. Interest and experience in the cybersecurity domain is beneficial.
Initial references:
[1] Matteo Boffa, Idilio Drago, Marco Mellia, Luca Vassio, Danilo Giordano, Rodolfo Valentim, and Zied Ben Houidi. Logpr´ecis: Unleashing language models for automated malicious log analysis: Pr´ecis: A concise summary of essential points, statements, or facts. Computers Security, 141:103805, 2024.
[2] Yiren Chen, Mengjiao Cui, Ding Wang, Yiyang Cao, Peian Yang, Bo Jiang, Zhigang Lu, and Baoxu Liu. A survey of large language models for cyber threat detection.Computers Security, 145:104016, 2024.
[3] Eider Iturbe, Oscar Llorente-Vazquez, Angel Rego, Erkuden Rios, and Nerea Toledo. Unleashing offensive artificial intelligence: Automated attack technique code generation. Computers Security, 147:104077, 2024.
[4] Abdechakour Mechri, Mohamed Amine Ferrag, and Merouane Debbah. Secureqwen: Leveraging llms for vulnerability detection in python codebases. Computers Security,148:104151, 2025.
[5] Shuang Tian, Tao Zhang, Jiqiang Liu, Jiacheng Wang, Xuangou Wu, Xiaoqiang Zhu, Ruichen Zhang, Weiting Zhang, Zhenhui Yuan, Shiwen Mao, et al. Exploring the role of large language models in cybersecurity: A systematic survey. arXiv preprint arXiv:2504.15622, 2025.
[6] Vaishali Vinay. The evolution of agentic ai in cybersecurity: From single llm reasoners to multi-agent systems and autonomous pipelines. arXiv preprint arXiv:2512.06659, 2025.
[7] Xiaoqing Wang, Yuanjing Tian, Keman Huang, and Bin Liang. Practically implementing an llm-supported collaborative vulnerability remediation process: A teambased approach. Computers Security, 148:104113, 2025.
[8] Jie Zhang, Haoyu Bu, Hui Wen, Yongji Liu, Haiqiang Fei, Rongrong Xi, Lun Li, Yun Yang, Hongsong Zhu, and Dan Meng. When llms meet cybersecurity: A systematic literature review. Cybersecurity, 8(1):55, 2025.
[9] Yongheng Zhang, Tingwen Du, Yunshan Ma, Xiang Wang, Yi Xie, Guozheng Yang, Yuliang Lu, and Ee-Chien Chang. Attackg+: Boosting attack graph construction with large language models. Computers Security, 150:104220, 2025.
16. Comparing Approaches for Agentic AI Systems Risk Management and Governance
Supervisors: Elmar Kiesling, Muhammad IkhsanBackground:
Recent developments in Generative Artificial Intelligence have fostered a shift from passive, chat-based assistants to Agentic AI systems. In such systems, AI agents possess a high degree of autonomy and adaptability, i.e., they can reason, break down complex goals into sub-tasks, use external tools (like APIs and databases), and operate over extended periods with minimal human intervention. However, whereas multi-agent generative systems have demonstrated large potential to tackle complex tasks effectively, their characteristics also give rise to unique and substantial risks. Classic AI risk frameworks designed for static models are poorly equipped to handle the dynamic, unpredictable, and emergent behaviors of autonomous agents. Consequently, academia, industry, and international regulatory bodies are developing specialized risk management and governance frameworks.
Research problem:
As a preliminary step, you will first survey and organize the wide field of Agentic AI Systems Risk Management and Governance (risk assessment, analysis, management; threat modeling; systems monitoring and control; governance etc.) and define a manageable scope.
The objective of the thesis is then to systematically, review, category and compare approaches to Agentic AI risk management and governance within that scope. By analyzing the strengths and limitations of different frameworks – e.g., by applying them on selected use cases (real-world, hypothetical, or derived from examples used in the literature) -- you will identify best practices and outstanding gaps in securing autonomous AI workflows.
Initial starting points for potential research questions include:
Risk landscape:
What unique classes of risks arise in the context of agentic AI?
Along which dimensions can they be organized and compared?
What aspects of Agentic AI systems affect how risks need to be identified, analyzed, and managed?
How do current governance approaches compare:
What scope do they cover and what classes of unique risks in the agentic space do they aim to tackle (on regulatory, organizational, and technical levels)?
Are the approaches genuinely new, or mostly adaptations of existing GenAI/LLM risk methods?
How do they relate to broader frameworks (NIST AI RMF, ISO/IEC 42001etc.)?
What are technical gaps where current management frameworks fail to provide adequate support for agentic workflows?
Required skills:
Strong interest in Agentic AI
Interest applying and evaluating methods on a technical level
Experience with agentic frameworks and programming agentic systems is not strictly necessary, but beneficial.
Initial references:
17. Leveraging AI Tools for Translating Lecture Videos in Higher Education
Supervisor: Michael FeursteinBackground: The rapid spread of publicly accessible generative AI services raises the question of how these tools can be integrated into higher education (Miao & Holmes, 2023). One application is the translation of video-based lecture content. As the number of international students grows, so does the demand for course material in both German and English. A case in point is the course “Grundlagen der Wirtschaftsinformatik” (Foundations of Information Systems): its content was originally designed in German and subsequently translated into English by hand, a tedious and time-consuming process. The course comprises more than nine hours of video, and experience shows that producing a single 45-minute video takes roughly four to five hours of studio time recording the material. This does not include time spent editing the video into a finished product. Recent advances in automatic speech recognition (Radford et al., 2023), translation and speech synthesis have produced automatic dubbing tools that promise to automate much of this pipeline. Research, however, points to weaknesses in naturalness, timing and translation accuracy (Brannon et al., 2023), even though first applications to educational videos report encouraging results (Kubota, 2026). If AI-based tools can reliably reduce this workload, they could substantially relieve lecturers.
Research Problem: This bachelor thesis analyzes the feasibility of using AI tools and toolchains to translate German-language lecture videos into English. It may address the following questions:
Tool landscape: Which AI tools for transcription, translation, voice synthesis and dubbing are currently available, and what are their affordances?
Progress and limitations: Where do current tools still fall short?
Toolchain design: What could an automated toolchain for translating videos look like?
Evaluation: How well does this toolchain perform on sample course material in terms of output quality and time saved compared to manual re-recording?
Legal and ethical aspects: What must be considered regarding data protection, copyright, voice cloning and transparency obligations considering the EU AI Act (Miao & Holmes, 2023)?
References
Brannon, W., Virkar, Y., & Thompson, B. (2023). Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing. Transactions of the Association for Computational Linguistics, 11, 419–435. https://doi.org/10.1162/tacl_a_00551
Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
Kubota, T. (2026). Lowering barriers to statistics learning by localizing English-language educational videos with multilingual AI dubbing. In H. Mori & Y. Asahi (Eds.), Human Interface and the Management of Information. HCII 2026 (LNCS 16706, pp. 100–116). Springer. https://doi.org/10.1007/978-3-032-29178-3_7
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning (PMLR 202, pp. 28492–28518). https://proceedings.mlr.press/v202/radford23a.html
Keywords: AI-based video translation, higher education, lecture videos, automatic dubbing, machine translation, speech synthesis
18. Teaching Management Information Systems in the Age of AI – Rethinking Education in Information Systems Curricula
Supervisor: Michael FeursteinBackground: Generative AI (GenAI) has moved from research labs into everyday business practice (Feuerriegel et al., 2024), and it is reshaping both how Management Information Systems (MIS) is taught and what needs to be taught. Tools such as ChatGPT, Claude or AI agents can now write SQL queries, draft entity-relationship and BPMN models, analyse business cases, produce code and write essays. They solve many of the assignments that traditionally form the core of IS courses. Van Slyke et al. (2023) argue that IS educators must respond by rethinking learning objectives, assignments and assessment rather than simply banning these tools. At the same time, evidence on learning effects is mixed: in computing education, GenAI can support learners but also mask gaps in understanding (Denny et al., 2024), and a large field experiment shows that unrestricted access can harm learning, while tutors with pedagogical guardrails largely mitigate this effect (Bastani et al., 2025). In addition, IS graduates are expected to select, implement and govern AI systems in organizations, which raises the question of whether recommendation for study programs in MIS (Beverungen, 2025) still reflect what employers need. Lecturers face a twofold challenge: using AI as a teaching tool and treating AI as a teaching subject.
Research Problem: This bachelor thesis investigates how teaching in MIS and IS curricula should be adapted to the widespread availability of GenAI. It may address the following questions:
State of practice: How are students and lecturers currently using GenAI in typical IS courses (e.g., data modelling, process management, ERP systems, programming, IT management)?
Learning effects: What does current empirical research say about the effects of GenAI on learning outcomes?
Course design: How could learning objectives, teaching activities and assessment in a selected MIS course be redesigned (e.g., AI-assisted case work, guardrailed AI tutors, oral or process-oriented assessment)?
Evaluation: How do lecturers and students assess the proposed concept, e.g., based on expert interviews or a student survey?
Integrity and policy: What must be considered regarding academic integrity, transparency of AI use, data protection and institutional AI guidelines?
References
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
Beverungen, D. (2025). Rahmenempfehlung für Studiengänge in Wirtschaftsinformatik an Hochschulen (G. für I. e.V, Ed.). Gesellschaft für Informatik e.V. https://doi.org/10.18420/rec2025_67
Denny, P., Prather, J., Becker, B. A., Finnie-Ansley, J., Hellas, A., Leinonen, J., Luxton-Reilly, A., Reeves, B. N., Santos, E. A., & Sarsa, S. (2024). Computing Education in the Era of Generative AI. Communications of the ACM, 67(2), 56–67. https://doi.org/10.1145/3624720
Feuerriegel, S., Hartmann, J., Janiesch, C., & Zschech, P. (2024). Generative AI. Business & Information Systems Engineering, 66(1), 111–126. https://doi.org/10.1007/s12599-023-00834-7
Van Slyke, C., Johnson, R., & Sarabadani, J. (2023). Generative Artificial Intelligence in Information Systems Education: Challenges, Consequences, and Responses. Communications of the Association for Information Systems, 53(1), 1–21. https://doi.org/10.17705/1CAIS.05301
Keywords: information systems education, management information systems, generative AI, curriculum design, assessment design, academic integrity
19. Automated Extraction of High-Quality Model Cards
Supervisor: Fajar J. EkaputraDesign and develop a tool that automatically collects model cards from platforms such as Hugging Face (other platforms such as Kaggle) until a target number of cards meets a target documentation quality. The result is a reusable dataset of well-documented model cards for selected tasks (e.g., question answering and/or image-to-text), plus statistics on how hard such cards are to find.
Background
Model cards document what a machine learning model does, how it was evaluated, and where it can fail. In practice, many are sparse or filled with template placeholders. Bhat et al. (CHI 2023) analyzed public model cards and found a substantial gap between what the model card proposal suggests and what is actually documented; they also contributed a rubric for evaluating documentation quality.
Researchers who want to study model documentation, or who need good examples, currently have to search by hand. Collecting cards that are actually complete is the bottleneck, and this thesis automates it.
Tasks
You design and implement a configurable pipeline. The user sets the task groups, the target number of cards (N) and the quality threshold (T); the tool then loops until both are met.
Retrieve candidate model cards for the chosen task groups and remove near-duplicates (forks and fine-tunes often copy their parent's card).
Apply cheap rule-based filters, such as missing sections, very short text and placeholder phrases like "More information needed".
Score the remaining cards against the Bhat et al. rubric and extract the answers to its questions, using a self-hosted open-source LLM (about 35B parameters).
Check the quality threshold, keep qualifying cards, and fetch more candidates until N is reached.
(Optional - Bonus): KG construction from the extracted knowledge
Report the process: how many cards were inspected, how many passed each stage, yield per iteration and runtime.
Research questions
RQ1: How accurately can a tiered pipeline (rules, then LLM) score and extract rubric items, compared with rules alone or an LLM alone, measured against a hand-labelled gold set?
RQ2: How many candidates must be processed to obtain N cards at threshold T, and what does it cost in time and LLM calls?
RQ3: How does documentation quality differ between task groups (e.g., text-based and vision-based)?
What you will learn and what we expect
You will gain hands-on experience with the Hugging Face Hub API, LLM-based structured information extraction, and the evaluation of automatic classifiers against manual annotation. You will also learn how to report an empirical pipeline reproducibly.
Expected background: solid knowledge on programming (e.g., Python) and basic familiarity with machine learning or NLP. No prior LLM deployment experience is needed; the model is already hosted.
Expected Outcomes
The working pipeline tool with documentation
A curated dataset of model cards meeting N and T, with extracted rubric answers
A gold-set evaluation of the scoring accuracy
A bachelor thesis report
Keywords
Model cards, Information extraction, LLMs
Starting references
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., Gebru, T. (2019). Model Cards for Model Reporting. FAT* '19. The original model card proposal.
Bhat, A., Coursey, A., Hu, G., Li, S., Nahar, N., Zhou, S., Kästner, C., Guo, J. L. C. (2023). Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and Traceability. CHI '23. Source of the documentation rubric.
Liang, W., Rajani, N., Yang, X., Ozoani, E., Wu, E., Chen, Y., Smith, D. S., Zou, J. (2024). Systematic analysis of 32,111 AI model cards characterizes documentation practice in AI. Nature Machine Intelligence, 6(7), 744-753. Large-scale analysis of Hugging Face model cards.
Hugging Face. Create and share Model Cards. huggingface_hub documentation, including ModelCard.load() for retrieving cards.
20. The Effects of Human-AI Collaboration on Long-Term Human Learning
Supervisors: Stefani Tsaneva, Marta SabouKeywords: Human-AI, Learning, Knowledge Transfer, AI Assistance
Context: Hybrid (Human-AI) intelligence refers to settings in which human actors and artificial intelligence (AI) systems collaborate to accomplish a task, combining human judgment with the speed and processing power of AI. A central principle of this approach is human augmentation: AI should complement and enhance human capabilities rather than replace them.
Much research has focused on understanding, designing, and optimizing hybrid intelligence teams for efficiency and output quality. However, less is known about how working with AI affects the human teammate, particularly in terms of long-term learning.
While the concept of co-learning- where humans learn from AI and AI learns from humans- has received considerable research attention, studies have largely focused on learning during the collaboration itself. The effects of human-AI teaming on the human’s ability to perform independently after the collaboration ends remain underexplored.
Problem: It remains unclear whether AI collaboration truly augments human capabilities or whether increased reliance on AI may instead hinder the development and exercise of those capabilities.
Goal/expected results of the thesis: Identify and synthesize existing research that examines the learning effects of human-AI collaboration on the human participant, with a particular focus on learning after the collaboration.
Research Questions (RQs):
What types of learning outcomes have been studied in human-AI collaboration?
What types of tasks have been selected to study learning outcomes of human-AI collaboration?
Which academic venues and research fields have published research on the learning effects of human-AI collaboration?
What evidence exists for positive or negative effects of AI collaboration on learning and independent performance?
What factors influence whether AI collaboration supports or hinders human learning?
To what extent are individual characteristics of the human participant (e.g., gender, educational background, qualifications, and profession) reported in studies examining the learning effects of human-AI collaboration?
Methodology: Literature review (reports for each step to be included)
Search Strategy: Define keywords, select search scope, select digital search libraries, perform search.
Paper Selection: Identify papers relevant to the research topic and select approximately 15–20 papers for detailed analysis.
Data Extraction: Extract relevant information (to address the RQs) from each selected paper.
Data Analysis: Analyse key trends, potential research gaps and limitations
Report: Present and synthesize the analysis results in relation to the RQs.
Required Skills:
Strong critical and analytical thinking
Basic data analysis skills
References
Shi, Q., Jimenez, C., Yao, S., Haber, N., Yang, D. and Narasimhan, K., 2026. When models know more than they can explain: Quantifying knowledge transfer in human-ai collaboration. Advances in Neural Information Processing Systems, 38, pp.105135-105170.
Mozannar, H., Chen, V., Alsobay, M., Das, S., Zhao, S., Wei, D., Nagireddy, M., Sattigeri, P., Talwalkar, A. and Sontag, D., 2025. The RealHumanEval: evaluating large language models' abilities to support programmers. Transactions on Machine Learning Research. https://openreview.net/pdf?id=hGaWq5Buj7
21. Sortition with diversity of opinion
Supervisors: Jan Maly, Felicia SchmidtKeywords: computational social choice, sortition, representation
Context: Democracies increasingly look for ways to involve citizens directly in political decisions, not only through elections. One popular approach is the citizens' assembly: a small group of ordinary people who discuss a political question in depth and give recommendations, for example on climate policy or local budgets. Since not everyone can take part, the central question is how to choose this group fairly.
The classic answer is sortition, selection by lottery. To make sure the group looks like the population, the lottery is usually constrained by "hard" criteria such as age, gender, region, or educational background. Computational social choice, the field that studies collective decision-making with algorithmic tools, has developed methods that do this while giving every person a fair chance of being picked.
However, matching the population on background characteristics does not guarantee that the group contains the full range of opinions held in society. Two people of the same age and background can think very differently about an issue. The opposite approach, discursive representation, aims to include the diversity of viewpoints directly. Here, people are asked a set of questions about the topic, and the committee is chosen so that the answers it contains are as varied as possible. This captures opinions well, but it can end up with a committee that is skewed in terms of age, background, or other demographics, and it gives no guarantee that everyone has a fair chance of selection.
Problem: Both approaches capture representation ideals, but currently there is no established method that combines the two: a selection procedure that is fair and representative in terms of background and covers a wide diversity of opinion.
Goal/expected results of the thesis: This thesis will develop and study an algorithm for selecting a committee that satisfies two goals at once:
Fairness: the selection is a lottery in which every person has a non-zero chance of being chosen.
Diversity of opinion: among the possible lotteries, the committees drawn should contain as much diversity of opinion as possible.
Research Questions:
How can a committee selection procedure be designed so that every person has a chance of being picked while the diversity of opinion of the resulting committees is as high as possible?
What is the trade-off between fairness (equal chances, matching demographic quotas) and diversity of opinion?
How does the proposed method perform compared to standard sortition and to purely opinion-based selection, on real or simulated data?
Methodology:
Literature review: Get familiar with computational social choice, in particular sortition and committee selection, and with the notion of discursive representation.
Formalization: Turn the informal goals (fairness, diversity of opinion) into precise measures that can be computed.
Design and implementation: Develop an algorithm that combines both goals and implement it, ideally in Python.
Evaluation: Test the algorithm on real or synthetic survey data and analyze the results.
Required Skills:
A good understanding of data analysis, ideally in Python.
Willingness to learn about mathematical measures of fairness and diversity. You don't need prior knowledge of computational social choice, but you should be comfortable reading and working with simple formal definitions.
Independent work with regular meetings with the supervisors: you present your progress, we give feedback and help you set the direction.
References
Flanigan, B., Gölz, P., Gupta, A. et al. Fair algorithms for selecting citizens’ assemblies. Nature 596, 548–552 (2021). https://doi.org/10.1038/s41586-021-03788-6
DRYZEK JS, NIEMEYER S. Discursive Representation. American Political Science Review. 2008;102(4):481-493. https://doi.org/10.1017/S0003055408080325
Soroush Ebadian, Gregory Kehne, Evi Micha, Ariel D. Procaccia, Nisarg Shah: Is Sortition Both Representative and Fair? NeurIPS 2022
22. Voting with Expressed Disapprovals and its Link to Pigouvian Tax
Supervisors: Hanna Kern, Jan MalyKeywords: Computational Social Choice, Pigouvian Tax, Literature Review, Fairness, Democracy, Voting, Simulations, Simulations in Python.
Context:
A growing number of novel digital participation processes allow citizens to express their opinions and directly influence policy decisions on a wide range of topics, from the design of a new park in Vienna [1] to the shape of the new constitution in Chile [2] and Iceland [3]. One major challenge in such digital democracy processes is the fair representation of minority opinions. Computer scientists have, in recent years, developed novel tools and algorithms that can be used to make the different forms of digital participation fairer and more representative [4].
In this thesis, we will focus on a setting that has not received a lot of attention so far, namely on elections where voters can express which candidates or opinions they approve of and which they disapprove of - a model that captures in particular many large scale deliberation processes, hosted on platforms like Pol.is. Kraiczy et al. (2025) [5] recently explored two settings with disapprovals and suggested fairness notions and voting rules, where the Asymmetric setting has a link to Pigouvian tax which is worth exploring further.
Problem:
Research into voting with approvals and disapprovals is very new and only theoretical, the paper by Kraiczy et al. (2025) proposes a very specific cost scheme without discussing any other potential tax functions.
Goal/expected results of the thesis:
This thesis will summarise available literature on Pigouvian tax and investigate the link between it and the setting proposed by Kraiczy et al. (2025), along with simulating the effect project cost has on the election outcome.
Research Questions:
What is known about Pigouvian tax and how to calculate it?
How do you convert the idea of Pigouvian tax to an election with expressed approvals and disapprovals?
How do different choices for the tax affect the outcome of (simulated) deliberation processes?
Methodology:
Get familiar with the setting and the intuition behind the Asymmetric setting proposed by Kraiczy et el. (2025).
Perform a literature review on Pigouvian tax.
Simulations in python.
Required Skills:
Good understanding of Pigouvian tax.
Ability to read, summarize, and explain available literature
Good understanding of data analysis, ideally with python.
A willingness to learn about mathematical measures of fairness.
References and related literature:
[1] https://mitgestalten.wien.gv.at/de-DE/projects/miep-gies-park
[2] https://europeandemocracyhub.epd.eu/wp-content/uploads/2023/12/Case-Study-Chile-FINAL-v2.pdf
[3] Hélène Landemore, When public participation matters: The 2010–2013 Icelandic constitutional process, International Journal of Constitutional Law, Volume 18, Issue 1, January 2020, Pages 179–205,https://doi.org/10.1093/icon/moaa004
[4] Lackner, Martin, and Piotr Skowron. Multi-winner voting with approval preferences. Springer Nature, 2023.
[5] Sections 1,2,4 of:
Kraiczy, Sonja, Georgios Papasotiropoulos, and Piotr Skowron. "Proportionality in Thumbs Up and Down Voting." arXiv preprint arXiv:2503.01985 (2025).
Mankiw, N. Gregory. Principles of Economics. 5th ed., South-Western Cengage Learning, 2011.
Boehmer, Niclas, et al. "Guide to numerical experiments on elections in computational social choice." arXiv preprint arXiv:2402.11765 (2024).
23. Prediction-Guided Preference Elicitation: Asking Voters the Right Questions
Supervisor: Jan MalyKeywords: Computational Social Choice, Preference Elicitation, Minimax Regret, Voting, Machine Learning, Prediction, Simulations in Python.
Context:
Digital tools allow us to make many more decisions in a democratic way by asking all stakeholders for their preferences. However, asking every participant to rank all options quickly becomes too much work, especially when there are many alternatives. Preference elicitation tackles this by asking participants only a few well-chosen questions (for example, "Do you prefer A to B?") and stopping once enough is known to pick a good outcome. Lu and Boutilier (2020) [2] proposed an influential framework for this. It uses minimax regret to measure how far the current best choice could be from the true winner, and it picks the next question according to how much that question could reduce this regret.
Problem:
The question-selection strategy of Lu and Boutilier (2020) is distribution-free. When it judges how useful a question is, it looks at the regret reduction under a single assumed (optimistic) answer and ignores how likely that answer actually is. In practice, answers are often quite predictable, for instance from the voter’s earlier answers, from other voters with similar preferences, or from a statistical model of the population. Questions whose answers we can already guess with high confidence tell us little. Other questions may look promising only because they assume an answer that is very unlikely.
Goal/expected results of the thesis:
The thesis will extend the minimax regret elicitation framework with a prediction step. Before a question is chosen, the probability of each possible answer is estimated, and questions are then ranked by the expected reduction in minimax regret. The student will design one or more ways to predict answers, ranging from simple frequency-based estimates to probabilistic models of rankings or standard machine learning methods. These strategies will be compared in simulations against the original approach on synthetic and real-world preference data.
Research Questions:
How can the probability of a voter’s answer to an elicitation query be predicted from the information available so far?
How can these predictions be built into the minimax-regret-based choice of the next query?
Does prediction-guided elicitation reach a low-regret or provably optimal outcome with fewer queries than the original strategy, and how does this depend on the data (e.g., how homogeneous voters’ preferences are)?
Methodology:
Get familiar with minimax regret and the elicitation strategy of Lu and Boutilier (2020), focusing on a single voting rule such as Borda.
Review basic approaches for predicting preferences (e.g., the Mallows model, frequency-based estimates, simple classifiers).
Implement the original strategy and the new prediction-guided variants in Python.
Run simulations on synthetic data and on real preference data (e.g., from PrefLib).
Required Skills:
Good programming skills in Python.
Basic knowledge of probability and statistics; some machine learning experience is a plus.
Ability to read and understand scientific papers.
Willingness to learn about voting rules and algorithmic decision-making.
References and related literature:
[1] Sections 1–5 of: Lu, Tyler, and Craig Boutilier. "Preference elicitation and robust winner determination for single- and multi-winner social choice." Artificial Intelligence 279 (2020): 103203. doi.org/10.1016/j.artint.2019.103203
Kalech, Meir, et al. "Practical voting rules with partial information." Autonomous Agents and Multi-Agent Systems 22.1 (2011): 151–182.
Mallows, Colin L. "Non-null ranking models. I." Biometrika 44.1/2 (1957): 114–130.
Mattei, Nicholas, and Toby Walsh. "PrefLib: A library for preferences." International Conference on Algorithmic Decision Theory (ADT), 2013.
Boehmer, Niclas, et al. "Guide to numerical experiments on elections in computational social choice." arXiv preprint arXiv:2402.11765 (2024).
24. Automated evaluation of AI-generated explanations
Supervisors: Stefani Tsaneva, Marta SabouKeywords: ontology engineering, human-centric explanations, large language models
Context: Knowledge Engineering (KE) encompasses a variety of activities, including the acquisition of knowledge and its representation through semantic models such as ontologies. Traditionally, KE requires substantial manual effort to define, implement, and validate domain-specific requirements. Moreover, tool support for many KE tasks remains limited, increasing the likelihood of modeling errors, especially when ontology engineers lack advanced KE training or are working with complex logical constraints. Recently, to support ontology engineers, the potential of Large Language Models (LLMs) has been explored in the context of ontology verification, specifically for defect detection, classification, explanation, and correction. While initial studies demonstrate that LLMs can assist with these tasks, further experimentation is necessary to generalize and extend these findings.
Problem: Currently, there is a lack of tools and methods for the automated evaluation of AI-generated explanations within the context of ontology validation.
Goal/expected results of the thesis: This thesis will investigate how LLMs can be utilised to annotate AI-generated explanations according to value-based requirements.
Research Questions: To what extent can LLMs annotate AI-generated explanations according to value-based requirements?
How accurately do LLMs evaluate ontology defect explanations?
Do different LLMs vary in their evaluation performance?
How consistent are LLM-generated annotations across repeated evaluations of the same input?
Methodology:
Get familiar with prior experiments on LLMs for ontology defect explanation and the produced explanations dataset.
Design and implement scripts (e.g., in Python or other suitable languages) to prompt various LLMs for explanation evaluation tasks.
Perform experiments across different LLMs and analyse the results.
Required Skills:
Understanding of ontologies, ontology constraints and reasoning (SBWL K2 completed).
Experience with Python (or other languages that support API access to LLMs).
Data analysis skills for processing generated outputs and evaluating performance.
References
Tsaneva, S., Herwanto, G. B., Llugiqi, M., & Sabou, M. Knowledge Engineering with Large Language Models: A Capability Assessment in Ontology Evaluation. https://www.semantic-web-journal.net/system/files/swj3852.pdf
C.-H. Chiang, H.-y. Lee, Can large language models be an alternative to human evaluations?, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 2023. 10.18653/v1/2023.acl-long.870
25. Profiling Knowledge Graphs: How Realistic are Benchmark KGs Compared to Real-World Knowledge Graphs?
Supervisors: Majlinda Llugiqi, Marta SabouKeywords: Knowledge Graphs, KG characteristics, Benchmark Datasets, Knowledge Graph Embeddings
Context: Knowledge graphs (KGs) represent entities and their relationships in a structured form and are used across many domains. Knowledge graph embedding (KGE) methods turn KGs into vector representations, which are used for tasks such as link prediction and as features in machine learning models. Research has shown that the characteristics of a KG, such as its node degree distribution, relation frequency or the presence of symmetric and inverse relations, influence how well embedding methods perform.
Problem: KGE methods are mostly developed and evaluated on a small set of benchmark KGs (e.g., FB15k-237, WN18RR). These benchmarks have been criticised as structurally unrealistic, yet little is known about how they actually differ from KGs used in practice. As a result, it is unclear whether results obtained on benchmarks transfer to real-world, domain-specific KGs.
Goal/expected results of the thesis:
The goal is to systematically measure and compare the characteristics of benchmark and real-world KGs. The expected results are:
an overview of KG characteristics relevant for KG embeddings and how to measure them;
a profiling tool in Python that computes these characteristics for a given KG;
a quantitative comparison of benchmark and real-world KGs, showing where they differ and which benchmarks best resemble real-world KGs.
Research Questions:
Which structural and semantic characteristics can be used to describe and compare KGs, and how can they be measured?
How do benchmark KGs differ from real-world, domain-specific KGs with respect to these characteristics?
Which benchmark KGs best resemble real-world KGs, and what does this imply for the evaluation of KG embedding methods?
Methodology:
Literature review to identify KG characteristics relevant for KG embeddings and existing ways to measure them.
Implementation of a profiling tool in Python (e.g., using NetworkX, RDFLib or PyKEEN).
Application of the tool to selected benchmark KGs (e.g., FB15k-237, WN18RR) and real-world KGs (e.g., domain-specific KGs found in the literature).
Comparative analysis and visualisation of the results.
Required Skills:
Programming skills in Python.
Basic understanding of knowledge graphs and graph concepts (e.g., nodes, edges, degree). Knowledge of KG embeddings is an advantage but not required.
Basic data analysis and visualisation skills.
References:
[1] Sardina, Jeffrey, John D. Kelleher, and Declan O'Sullivan. "A Survey on Knowledge Graph Structure and Knowledge Graph Embeddings." arXiv preprint arXiv:2412.10092 (2024).
[2] Rossi, Andrea, and Antonio Matinata. "Knowledge graph embeddings: Are relation-learning models learning relations?." EDBT/ICDT Workshops. Vol. 2578. 2020.
[3] Akrami, Farahnaz, Mohammed Samiul Saeef, Qingheng Zhang, Wei Hu, and Chengkai Li. "Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental Study." Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data (2020).
26. Scientific Recommendations from KGs using GraphRAGs and Agents
Supervisor: Diego Alberto Rincon YanezBackground
Scientific collaboration involves identifying relevant researchers, publications, and relationships for a research problem, but scholarly data is often spread across diverse sources with heterogeneous metadata. Knowledge Graphs (KGs) are structured representations of knowledge in which real-world entities are modeled as nodes and their relationships as explicit, machine-readable links. Unlike conventional tabular databases, KGs are well suited to integrating heterogeneous data from multiple sources while preserving the meaning and context of relationships between entities. In the scholarly domain, this is especially valuable because scientific information is naturally interconnected through researchers, publications, institutions, venues, topics, citations, projects, and collaborations. These elements are represented in a Scholarly Knowledge Graph. GraphRAG and agent-based systems can provide an intelligent access layer over the resulting Scholarly Knowledge Graph. GraphRAG retrieves relevant subgraphs to ground answers in scholarly evidence, while agents with specialized skills can query publications, trace citation paths, analyze co-authorship networks, compare researcher profiles, and combine information from multiple graph regions. This enables more exploratory discovery, such as identifying hidden collaboration opportunities, finding researchers connected through complementary expertise, or revealing indirect relationships between topics, publications, and scientific communities.
Research Problem
Metadata extracted from files and other resources alone provides only a partial representation of the scholarly environment. Additional information, including persistent identifiers, complete author information, affiliations, venues, references, and related scholarly entities, must be retrieved from external sources, reconciled, and represented consistently. The project gathers essential data, including researchers, publications, authorship, venues, identifiers, keywords, and affiliations, from publicly accessible scientific databases through API requests. The project therefore addresses the following question: How can scholarly data and metadata be integrated into a knowledge graph (KG) and subsequently utilized through artificial intelligence (AI) methods to facilitate the discovery of relevant researchers, publications, collaboration opportunities, and potential scientific collaborators?
Initial References
Dessí, D., Osborne, F., Reforgiato Recupero, D., Buscaldi, D., & Motta, E. (2022). CS-KG: A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics): 13489 LNCS (pp. 678–696). doi.org/10.1007/978-3-031-19433-7_39.
Ramírez Molina, A. A., Li, S., & Gong, J. (2026). Mapping scholarly knowledge: A systematic review of Knowledge Graphs for academic papers. Information Processing and Management, 63(7), 104842. doi.org/10.1016/j.ipm.2026.104842.
Jaradeh, M. Y., Singh, K., Stocker, M., & Auer, S. (2021). Triple Classification for Scholarly Knowledge Graph Completion. K-CAP 2021 - Proceedings of the 11th Knowledge Capture Conference, 225–232. doi.org/10.1145/3460210.3493582.
Keywords
Knowledge Graphs; Scholarly Knowledge Graphs; Semantic Web; Scholarly Metadata; Data Integration.
Prior Knowledge
Python programming skills
Basic Graph Analytics
Familiarity with REST APIs is desirable.
Usage of LLMs with Python
Write a Thesis