Dr. David Broneske studierte Informatik im Bachelor und Master an der Otto-von-Guericke Universität Magdeburg, wo er 2019 promovierte und darauf seine Habilitation begann. Von 2019 bis 2020 übernahm er die Vertretung des Lehrstuhls Datenbank und Informationssysteme an der Hochschule Anhalt in Köthen. Er hat im März 2021 die kommissarische Leitung der Abteilung 4

Dr. David Broneske
Abteilung Infrastruktur und Methoden
Abteilungsleitung
- 0511 450670-454
- Google Scholar
- Orcid
Wissenschaftliche Forschungsgebiete
Forschungsdatenmanagement für Learning Analytics Daten, Hauptspeicherdatenbanken auf moderner Hardware, Interaktive Datenexploration und Visualisierung, Künstliche Intelligenz für Datenbereinigung und Datenanalyse
Liste der Projekte
Liste der Publikationen
The imitation game: Evaluating persona-driven LLM response behavior in web surveys.Shahania, S., Spiliopoulou, M., & Broneske, D. (2026).The imitation game: Evaluating persona-driven LLM response behavior in web surveys. In Wong, R. et al. (Hrsg.), Advances in Knowledge Discovery and Data Mining (S. 474-486). Singapore: Springer. https://doi.org/10.1007/978-981-92-1468-6_33 |
GraphAware: Interpretable machine learning on graphs.Walke, D., Steinbach, D., Schönhuth, A., Saake, G., Broneske, D., & Heyer, R. (2026).GraphAware: Interpretable machine learning on graphs. Discover Artificial Intelligence(1), 1-345. https://doi.org/10.1007/s44163-026-01028-2 |
MorphCraft: Morphology-aware transformers for Armenian and Greek.Avetisyan, H., Karasavva, C., & Broneske, D. (2026).MorphCraft: Morphology-aware transformers for Armenian and Greek. In IEEE Institute of Electrical and Electronic Engineers (Hrsg.), IEEE Xplore (S. 1-6). Jacksonville, Florida, USA: IEEE Xplore. |
Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification.Safikhani, P., & Broneske, D. (2025).Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification. ACL Anthology, 1018-1027. Abstract
While Automated Machine Learning (AutoML) systems have shown strong performance on structured data, their application to natural language processing (NLP) tasks remains limited by static, task-agnostic search spaces. In this work, we propose a context-aware extension of AutoPyTorch that dynamically adapts both the hyperparameter search space and neural architecture configuration based on corpus-level meta-features. Our approach extracts interpretable textual statistics—such as average sequence length, vocabulary richness, and class imbalance—to guide the configuration of key hyperparameters. We also introduce two adaptive neural backbones, whose structures are shaped by these meta-features to improve model expressiveness and generalization. |
SimKit: Similarity graphs, eigendecomposition and spectral clustering in Neo4j.Mondal, R., Ignatova, E., Heinzmann, J., Do, M. D., Murali, A., ... & Heyer, G. (2025).SimKit: Similarity graphs, eigendecomposition and spectral clustering in Neo4j. In IEEE Institute of Electrical and Electronic Engineers (Hrsg.), IEEE xplore (S. 685-691). Jacksonville, Florida, USA: IEEE Xplore. |
Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in potein-protein interaction networks.Abbas, M., Broneske, D., & Saake, G. (2025).Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in potein-protein interaction networks. In Arai, K. (Hrsg.), Intelligent Systems and Applications. Proceedings of the 2025 Intelligent Systems Conference (IntelliSys) (S. 512-528). Cham: Springer. https://doi.org/10.1007/978-3-031-99958-1_32 |
Static and dynamic contextual embedding for AutoML in text classification tasks.Safikhani, P., & Broneske, D. (2025).Static and dynamic contextual embedding for AutoML in text classification tasks. In IEEE Institute of Electrical and Electronic Engineers (Hrsg.), 2025 7th International Conference on Natural Language Processing (ICNLP) (S. 292-301). Jacksonville, Florida, USA: IEEE Xplore. https://doi.org/10.1109/ICNLP65360.2025.11108687 |
Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys.Shahania, S., Spiliopoulou, M., & Broneske, D. (2025).Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys. In Association for Computing Machinery (Hrsg.), CI '25: Proceedings of the ACM Collective Intelligence Conference (S. 113-128). New York: ACM. https://doi.org/10.1145/3715928.3737491 |
SBC-SHAP: Increasing the accessibility and interpretability of machine learning algorithms for sepsis prediction.Walke, D., Steinbach, D., Kaiser, T., Schönhuth, A., Saake, G., Broneske, D., & Heyer, R. (2025).SBC-SHAP: Increasing the accessibility and interpretability of machine learning algorithms for sepsis prediction. The Journal of Applied Laboratory Medicine. https://doi.org/10.1093/jalm/jfaf091 |
Edges are all you need: Potential of medical time series analysis on complete blood count data with graph neural networks.Walke, D., Steinbach, D., Gibb, S., Kaiser, T., Saake, G., ... & Heyer, R. (2025).Edges are all you need: Potential of medical time series analysis on complete blood count data with graph neural networks. PLOS One. https://doi.org/10.1371/journal.pone.0327636 |
SurveyBot: A new era of web survey pretesting.Shahania, S., Spiliopoulou, M., & Broneske, D. (2025).SurveyBot: A new era of web survey pretesting. In I. Maglogiannis, L. Iliadis, A. Andreou, & A. Papaleonidas (Hrsg.), Artificial Intelligence Applications and Innovations. AIAI 2025. IFIP Advances in Information and Communication Technology. Cham: Springer. https://doi.org/10.1007/978-3-031-96235-6_29 |
Towards automatic bias analysis in multimedia journalism.Hinrichs, R., Steffen, H., Avetisyan, H., Broneske, D., & Ostermann, J. (2025).Towards automatic bias analysis in multimedia journalism. Discover Artificial Intelligence, 5(1), 1-28. https://doi.org/10.1007/s44163-025-00362-1 |
Embracing NVM: Optimizing $B^𝜖$-tree structures and data compression in storage engines.Karim, S., Wünsche, F., Broneske, D., Kuhn, M., & Saake, G. (2025).Embracing NVM: Optimizing $B^𝜖$-tree structures and data compression in storage engines. In Binnig, C. et al. (Hrsg.), Datenbanksysteme für Business, Technologie und Web - Workshopband (BTW 2025) (S. 329-333). Bonn: Gesellschaft für Informatik. https://doi.org/10.18420/BTW2025-137 |
A multi-objective evolutionary algorithm for detecting protein complexes in PPI networks using gene ontology.Abbas, M. N., Broneske, D., & Saake, G. (2025).A multi-objective evolutionary algorithm for detecting protein complexes in PPI networks using gene ontology. Scientific Reports, 15. https://doi.org/10.1038/s41598-025-01667-y |
AutoML meets hugging face: Domain-aware pretrained model selection for text classification.Safikhani, P., & Broneske, D. (2025).AutoML meets hugging face: Domain-aware pretrained model selection for text classification. In A. Ebrahimi, S. Haider, E. Liu, M. L. Pacheco, & S. Wein (Hrsg.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop). Albuquerque, USA: Association for Computational Linguistics. Abstract
The effectiveness of embedding methods is crucial for optimizing text classification performance in Automated Machine Learning (AutoML). However, selecting the most suitable pre-trained model for a given task remains challenging. This study introduces the Corpus-Driven Domain Mapping (CDDM) pipeline, which utilizes a domain-annotated corpus of pre-fine-tuned models from the Hugging Face Model Hub to improve model selection. Integrating these models into AutoML systems significantly boosts classification performance across multiple datasets compared to baseline methods. Despite some domain recognition inaccuracies, results demonstrate CDDM’s potential to enhance model selection, streamline AutoML workflows, and reduce computational costs. |
Liste der Vorträge & Tagungen
Evaluating the effectiveness of CAPTCHAs and attention checks for LLM-driven survey bots.Shahania, S., Niemann-Lenz, J., & Broneske, D. (2026, Juni).Evaluating the effectiveness of CAPTCHAs and attention checks for LLM-driven survey bots. Vortrag auf der Konferenz 16th Scientific Conference on Data Collections Methods hosted by the ADM, the ASI and the Federal Statistical Office, Wiesbaden, Germany. |
Lexikalisch verloren, semantisch gefunden: LLM-gestützte Suche auf der Grundlage von DDI-Metadaten.Broneske, D., Daniel, A., Buck, D., & Weber, A. (2026, Juni).Lexikalisch verloren, semantisch gefunden: LLM-gestützte Suche auf der Grundlage von DDI-Metadaten. Vortrag im Rahmen der 10. Konferenz für Sozial- und Wirtschaftsdaten (10|KSWD), Berlin. Abstract
Damit Forschungsdaten für Forschende praktisch auffindbar sind, braucht es Suchsysteme, die Daten entlang konkreter Forschungsfragen zugänglich machen. Forschungsdatenzentren machen ihre Metadaten derzeit häufig mithilfe lexikalischer Suchmaschinen (z. B. Elasticsearch) durchsuchbar und liefern Ergebnisse auf Basis lexikalischer Übereinstimmungen. Das ermöglicht eine Volltextsuche und ein elaboriertes Relevanzranking der Suchergebnisse, hat jedoch den Nachteil, dass es keine semantisch relevanten Metadaten findet, wenn fachspezifische Konzepte oder Forschungsfragen nicht ausdrücklich in den Metadaten benannt sind. [...] Vollständiger Abstract: https://zenodo.org/records/20844070 |
The imitation game: Evaluating persona-driven LLM response behavior in web surveys.Shahania, S. (2026, Juni).The imitation game: Evaluating persona-driven LLM response behavior in web surveys. Vortrag auf der Konferenz The 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), Hongkong, SAR, China. https://doi.org/10.1007/978-981-92-1468-6_33 |
Werkstattbericht - Ein KI-Tool für die automatisierte Verkodung von Berufsangaben mittels Open-Source-LLMs.Broneske, D. (2026, April).Werkstattbericht - Ein KI-Tool für die automatisierte Verkodung von Berufsangaben mittels Open-Source-LLMs. Vortrag im Rahmen des Austausches des Verbund Forschungsdaten Bildung (VerbundFDB), DIPF Frankfurt. |
Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification.Safikhani, P., & Broneske, D. (2025, September).Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification. Poster auf der Konferenz Italian Conference on Computational Linguistics (CLiC-it), Cagliari, Italien. |
Das BMFTR-Datenportal - Arbeitsstand der Weiterentwicklung und Präsentation der Testseiten sowie DASTAT - Umsetzung der künftigen Erhebungen.Grützmacher, J., & Osterburg, M. (2025, September).Das BMFTR-Datenportal - Arbeitsstand der Weiterentwicklung und Präsentation der Testseiten sowie DASTAT - Umsetzung der künftigen Erhebungen. Vortrag im Rahmen des Jahresarbeitsgesprächs der Projektgruppe "Informationssysteme", Referat S 22 im BMFTR, Berlin, Deutschland. |
(Offene) Forschungssoftware: Ein unverzichtbarer Pfeiler der modernen Wissenschaft?Heußner, A., Fritzsch, B., Broneske, D., Trilcke, P., & Lamprecht, A.-L. (2025, September).Teilnahme an der Podiumsdiskussion (Offene) Forschungssoftware: Ein unverzichtbarer Pfeiler der modernen Wissenschaft? auf der Tagung Informatik Festival 2025, Gesellschaft für Informatik, Hasso-Plattner-Institut, Universität Potsdam, Potsdam. |
Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in protein-protein interaction networks.Abbas, M., Broneske, D., & Saake, G. (2025, August).Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in protein-protein interaction networks. Vortrag auf der Konferenz IntelliSys 2025 : 11th Intelligent Systems Conference 2025, Amsterdam. |
Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys.Shahania, S. (2025, August).Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys. Vortrag auf der Konferenz ACM Collective Intelligence 2025, LA Jolle, San Diego, California, USA. https://doi.org/10.1145/3715928.3737491 |
AutoML meets hugging face: Domain-aware pretrained model selection for text classification.Safikhani, P. (2025, April/Mai).AutoML meets hugging face: Domain-aware pretrained model selection for text classification. Poster auf der Konferenz 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Albuquerque, USA. Abstract
The effectiveness of embedding methods is crucial for optimizing text classification performance in Automated Machine Learning (AutoML). However, selecting the most suitable pre-trained model for a given task remains challenging. This study introduces the Corpus-Driven Domain Mapping (CDDM) pipeline, which utilizes a domain-annotated corpus of pre-fine-tuned models from the Hugging Face Model Hub to improve model selection. Integrating these models into AutoML systems significantly boosts classification performance across multiple datasets compared to baseline methods. Despite some domain recognition inaccuracies, results demonstrate CDDM’s potential to enhance model selection, streamline AutoML workflows, and reduce computational costs. |
Static and dynamic contextual embedding for AutoML in text classification tasks.Safikhani, P. (2025, März).Static and dynamic contextual embedding for AutoML in text classification tasks. Vortrag auf der Konferenz International Conference on Natural Language Processing (ICNLP 2025), Guangzhou, China. |
Enhancing AutoML for NLP: Context-aware hyperparameter tuning and text representation using large language models.Safikhani, P. (2025, Februar).Enhancing AutoML for NLP: Context-aware hyperparameter tuning and text representation using large language models. Vortrag auf dem Kolloquium Doktorandentag at Otto von Guericke University Magdeburg, Faculty of Computer Science, Magdeburg, Germany. |
Seit 03/2021
Kommissarischer Leiter der Abteilung 4
- SoSe 2016, 2020 & 2021 Advanced Topics in Databases (OVGU; ca. 50 Teilnehmer) - Vorlesung
- WiSe 2020 Data Warehouse Technologies (OVGU; ca. 80 Teilnehmer) - Vorlesung & Übung
- WiSe 2020 Datenbanken I (OVGU; ca. 270 Teilnehmer) - Vorlesung & Übung
- SoSe 2013 & 2020 Datenbanken Implementierungstechniken (OVGU; ca. 40 Teilnehmer) - Übung
- WiSe 2019 Moderne Datenbankkonzepte (HS-Anhalt; ca. 12 Teilnehmer) - Vorlesung & Übung
- WiSe 2019 Datenbanksysteme (HS-Anhalt; ca. 40 Teilnehmer ) - Vorlesung & Übung
- SoSe 2015-2019 Database Concepts (OVGU; ca. 120 Teilnehmer) - Übung
- WiSe 2012-2018 Datenbanken I (OVGU, ca. 270 Teilnehmer) - Übung
- SoSe 2012-2014 Datenmanagement (OVGU; ca. 100 Teilnehmer) - Übung
- 2020 - Distinguished Reviewer at Information Systems, Elsevier
- 2017 - Forschungspreis der Fakultät für Informatik der Otto-von-Guericke-Universität Magdeburg