David Broneske

Dr. David Broneske

Research Area Research Infrastructure and Methods
Head of Department
  • +49 511 450670-454
  • Google Scholar
  • Orcid

Dr. David Broneske did his Bachelor and Master in Computer Science at the Otto-von-Guericke University Magdeburg, where he received his Ph.D. in 2019 and afterward started his Habilitation. From 2019 to 2020, he was the substitutional head of the chair of Database and Informationssystems at Hochschule Anhalt in Köthen. Since March 2021, he is the acting head of Department 4

Read more Read less

Academic research fields

Research Data Management for Learning Analytics Data, Main-Memory Database Systems on Modern Hardware, Interactive Data Exploration and Visualization, Artificial Intelligence for Data Cleaning and Analysis

Projects

List of projects

Unfortunately, there is no result available for this search combination
The Appointment of Professors at Private and State Universities of Applied Sciences
Research cluster: Open Science
Publications

List of publications

Unfortunately, there is no result available for this search combination

The imitation game: Evaluating persona-driven LLM response behavior in web surveys.

Shahania, S., Spiliopoulou, M., & Broneske, D. (2026).
The imitation game: Evaluating persona-driven LLM response behavior in web surveys. In Wong, R. et al. (Hrsg.), Advances in Knowledge Discovery and Data Mining (S. 474-486). Singapore: Springer. https://doi.org/10.1007/978-981-92-1468-6_33

GraphAware: Interpretable machine learning on graphs.

Walke, D., Steinbach, D., Schönhuth, A., Saake, G., Broneske, D., & Heyer, R. (2026).
GraphAware: Interpretable machine learning on graphs. Discover Artificial Intelligence(1), 1-345. https://doi.org/10.1007/s44163-026-01028-2

MorphCraft: Morphology-aware transformers for Armenian and Greek.

Avetisyan, H., Karasavva, C., & Broneske, D. (2026).
MorphCraft: Morphology-aware transformers for Armenian and Greek. In IEEE Institute of Electrical and Electronic Engineers (Hrsg.), IEEE Xplore (S. 1-6). Jacksonville, Florida, USA: IEEE Xplore.

Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification.

Safikhani, P., & Broneske, D. (2025).
Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification. ACL Anthology, 1018-1027.
Abstract

While Automated Machine Learning (AutoML) systems have shown strong performance on structured data, their application to natural language processing (NLP) tasks remains limited by static, task-agnostic search spaces. In this work, we propose a context-aware extension of AutoPyTorch that dynamically adapts both the hyperparameter search space and neural architecture configuration based on corpus-level meta-features. Our approach extracts interpretable textual statistics—such as average sequence length, vocabulary richness, and class imbalance—to guide the configuration of key hyperparameters. We also introduce two adaptive neural backbones, whose structures are shaped by these meta-features to improve model expressiveness and generalization.

SimKit: Similarity graphs, eigendecomposition and spectral clustering in Neo4j.

Mondal, R., Ignatova, E., Heinzmann, J., Do, M. D., Murali, A., ... & Heyer, G. (2025).
SimKit: Similarity graphs, eigendecomposition and spectral clustering in Neo4j. In IEEE Institute of Electrical and Electronic Engineers (Hrsg.), IEEE xplore (S. 685-691). Jacksonville, Florida, USA: IEEE Xplore.

Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in potein-protein interaction networks.

Abbas, M., Broneske, D., & Saake, G. (2025).
Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in potein-protein interaction networks. In Arai, K. (Hrsg.), Intelligent Systems and Applications. Proceedings of the 2025 Intelligent Systems Conference (IntelliSys) (S. 512-528). Cham: Springer. https://doi.org/10.1007/978-3-031-99958-1_32

Static and dynamic contextual embedding for AutoML in text classification tasks.

Safikhani, P., & Broneske, D. (2025).
Static and dynamic contextual embedding for AutoML in text classification tasks. In IEEE Institute of Electrical and Electronic Engineers (Hrsg.), 2025 7th International Conference on Natural Language Processing (ICNLP) (S. 292-301). Jacksonville, Florida, USA: IEEE Xplore. https://doi.org/10.1109/ICNLP65360.2025.11108687

Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys.

Shahania, S., Spiliopoulou, M., & Broneske, D. (2025).
Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys. In Association for Computing Machinery (Hrsg.), CI '25: Proceedings of the ACM Collective Intelligence Conference (S. 113-128). New York: ACM. https://doi.org/10.1145/3715928.3737491

SBC-SHAP: Increasing the accessibility and interpretability of machine learning algorithms for sepsis prediction.

Walke, D., Steinbach, D., Kaiser, T., Schönhuth, A., Saake, G., Broneske, D., & Heyer, R. (2025).
SBC-SHAP: Increasing the accessibility and interpretability of machine learning algorithms for sepsis prediction. The Journal of Applied Laboratory Medicine. https://doi.org/10.1093/jalm/jfaf091

Edges are all you need: Potential of medical time series analysis on complete blood count data with graph neural networks.

Walke, D., Steinbach, D., Gibb, S., Kaiser, T., Saake, G., ... & Heyer, R. (2025).
Edges are all you need: Potential of medical time series analysis on complete blood count data with graph neural networks. PLOS One. https://doi.org/10.1371/journal.pone.0327636

SurveyBot: A new era of web survey pretesting.

Shahania, S., Spiliopoulou, M., & Broneske, D. (2025).
SurveyBot: A new era of web survey pretesting. In I. Maglogiannis, L. Iliadis, A. Andreou, & A. Papaleonidas (Hrsg.), Artificial Intelligence Applications and Innovations. AIAI 2025. IFIP Advances in Information and Communication Technology. Cham: Springer. https://doi.org/10.1007/978-3-031-96235-6_29

Towards automatic bias analysis in multimedia journalism.

Hinrichs, R., Steffen, H., Avetisyan, H., Broneske, D., & Ostermann, J. (2025).
Towards automatic bias analysis in multimedia journalism. Discover Artificial Intelligence, 5(1), 1-28. https://doi.org/10.1007/s44163-025-00362-1

Embracing NVM: Optimizing $B^𝜖$-tree structures and data compression in storage engines.

Karim, S., Wünsche, F., Broneske, D., Kuhn, M., & Saake, G. (2025).
Embracing NVM: Optimizing $B^𝜖$-tree structures and data compression in storage engines. In Binnig, C. et al. (Hrsg.), Datenbanksysteme für Business, Technologie und Web - Workshopband (BTW 2025) (S. 329-333). Bonn: Gesellschaft für Informatik. https://doi.org/10.18420/BTW2025-137

A multi-objective evolutionary algorithm for detecting protein complexes in PPI networks using gene ontology.

Abbas, M. N., Broneske, D., & Saake, G. (2025).
A multi-objective evolutionary algorithm for detecting protein complexes in PPI networks using gene ontology. Scientific Reports, 15. https://doi.org/10.1038/s41598-025-01667-y

AutoML meets hugging face: Domain-aware pretrained model selection for text classification.

Safikhani, P., & Broneske, D. (2025).
AutoML meets hugging face: Domain-aware pretrained model selection for text classification. In A. Ebrahimi, S. Haider, E. Liu, M. L. Pacheco, & S. Wein (Hrsg.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop). Albuquerque, USA: Association for Computational Linguistics.
Abstract

The effectiveness of embedding methods is crucial for optimizing text classification performance in Automated Machine Learning (AutoML). However, selecting the most suitable pre-trained model for a given task remains challenging. This study introduces the Corpus-Driven Domain Mapping (CDDM) pipeline, which utilizes a domain-annotated corpus of pre-fine-tuned models from the Hugging Face Model Hub to improve model selection. Integrating these models into AutoML systems significantly boosts classification performance across multiple datasets compared to baseline methods. Despite some domain recognition inaccuracies, results demonstrate CDDM’s potential to enhance model selection, streamline AutoML workflows, and reduce computational costs.

Presentations

List of presentations & conferences

Unfortunately, there is no result available for this search combination

Evaluating the effectiveness of CAPTCHAs and attention checks for LLM-driven survey bots.

Shahania, S., Niemann-Lenz, J., & Broneske, D. (2026, Juni).
Evaluating the effectiveness of CAPTCHAs and attention checks for LLM-driven survey bots. Vortrag auf der Konferenz 16th Scientific Conference on Data Collections Methods hosted by the ADM, the ASI and the Federal Statistical Office, Wiesbaden, Germany.

Lexikalisch verloren, semantisch gefunden: LLM-gestützte Suche auf der Grundlage von DDI-Metadaten.

Broneske, D., Daniel, A., Buck, D., & Weber, A. (2026, Juni).
Lexikalisch verloren, semantisch gefunden: LLM-gestützte Suche auf der Grundlage von DDI-Metadaten. Vortrag im Rahmen der 10. Konferenz für Sozial- und Wirtschaftsdaten (10|KSWD), Berlin.
Abstract

To ensure that research data is easily discoverable by researchers, search systems are needed that make data accessible in relation to specific research questions. Research data centres currently often make their metadata searchable using lexical search engines (e.g. Elasticsearch) and return results based on lexical matches. Whilst this enables full-text search and a sophisticated ranking of search results by relevance, it has the disadvantage of failing to identify semantically relevant metadata if subject-specific concepts or research questions are not explicitly named in the metadata. This is a problem, as theoretical constructs often exist only as latent dimensions [...] Full Abstract: https://zenodo.org/records/20844070

The imitation game: Evaluating persona-driven LLM response behavior in web surveys.

Shahania, S. (2026, Juni).
The imitation game: Evaluating persona-driven LLM response behavior in web surveys. Vortrag auf der Konferenz The 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), Hongkong, SAR, China. https://doi.org/10.1007/978-981-92-1468-6_33

Werkstattbericht - Ein KI-Tool für die automatisierte Verkodung von Berufsangaben mittels Open-Source-LLMs.

Broneske, D. (2026, April).
Werkstattbericht - Ein KI-Tool für die automatisierte Verkodung von Berufsangaben mittels Open-Source-LLMs. Vortrag im Rahmen des Austausches des Verbund Forschungsdaten Bildung (VerbundFDB), DIPF Frankfurt.

Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification.

Safikhani, P., & Broneske, D. (2025, September).
Context-aware search space adaptation of hyperparameters and architectures for AutoML in text classification. Poster auf der Konferenz Italian Conference on Computational Linguistics (CLiC-it), Cagliari, Italien.

Das BMFTR-Datenportal - Arbeitsstand der Weiterentwicklung und Präsentation der Testseiten sowie DASTAT - Umsetzung der künftigen Erhebungen.

Grützmacher, J., & Osterburg, M. (2025, September).
Das BMFTR-Datenportal - Arbeitsstand der Weiterentwicklung und Präsentation der Testseiten sowie DASTAT - Umsetzung der künftigen Erhebungen. Vortrag im Rahmen des Jahresarbeitsgesprächs der Projektgruppe "Informationssysteme", Referat S 22 im BMFTR, Berlin, Deutschland.

(Offene) Forschungssoftware: Ein unverzichtbarer Pfeiler der modernen Wissenschaft?

Heußner, A., Fritzsch, B., Broneske, D., Trilcke, P., & Lamprecht, A.-L. (2025, September).
Teilnahme an der Podiumsdiskussion (Offene) Forschungssoftware: Ein unverzichtbarer Pfeiler der modernen Wissenschaft? auf der Tagung Informatik Festival 2025, Gesellschaft für Informatik, Hasso-Plattner-Institut, Universität Potsdam, Potsdam.

Semantische Suche - Am Beispiel des BMFTR-Datenportal.

Broneske, D. (2025, September).
Semantische Suche - Am Beispiel des BMFTR-Datenportal. Vortrag im KI-Kolloquium am DZHW, Hannover.

Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in protein-protein interaction networks.

Abbas, M., Broneske, D., & Saake, G. (2025, August).
Improving the performance of evolutionary-based complex detection models using gene ontology-based mutation operator in protein-protein interaction networks. Vortrag auf der Konferenz IntelliSys 2025 : 11th Intelligent Systems Conference 2025, Amsterdam.

Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys.

Shahania, S. (2025, August).
Gotta catch 'em all... Or not?: How LLMs bypass traditional checks & mimic human response behavior in web surveys. Vortrag auf der Konferenz ACM Collective Intelligence 2025, LA Jolle, San Diego, California, USA. https://doi.org/10.1145/3715928.3737491

Herausforderungen der Datenanalyse in den Sozialwissenschaften am Beispiel des DZHW.

Broneske, D. (2025, Mai).
Herausforderungen der Datenanalyse in den Sozialwissenschaften am Beispiel des DZHW. Vortrag auf dem Kolloquium Advances in Database Research, Universität Salzburg, Salzburg.

AutoML meets hugging face: Domain-aware pretrained model selection for text classification.

Safikhani, P. (2025, April/Mai).
AutoML meets hugging face: Domain-aware pretrained model selection for text classification. Poster auf der Konferenz 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Albuquerque, USA.
Abstract

The effectiveness of embedding methods is crucial for optimizing text classification performance in Automated Machine Learning (AutoML). However, selecting the most suitable pre-trained model for a given task remains challenging. This study introduces the Corpus-Driven Domain Mapping (CDDM) pipeline, which utilizes a domain-annotated corpus of pre-fine-tuned models from the Hugging Face Model Hub to improve model selection. Integrating these models into AutoML systems significantly boosts classification performance across multiple datasets compared to baseline methods. Despite some domain recognition inaccuracies, results demonstrate CDDM’s potential to enhance model selection, streamline AutoML workflows, and reduce computational costs.

Bots in web survey interviews: A showcase.

Höhne, J. K., Claaßen, J., Shahania, S., & Broneske, D. (2025, März/April).
Bots in web survey interviews: A showcase. Vortrag im Rahmen der General Online Research (GOR) Conference, Berlin.

Static and dynamic contextual embedding for AutoML in text classification tasks.

Safikhani, P. (2025, März).
Static and dynamic contextual embedding for AutoML in text classification tasks. Vortrag auf der Konferenz International Conference on Natural Language Processing (ICNLP 2025), Guangzhou, China.

Enhancing AutoML for NLP: Context-aware hyperparameter tuning and text representation using large language models.

Safikhani, P. (2025, Februar).
Enhancing AutoML for NLP: Context-aware hyperparameter tuning and text representation using large language models. Vortrag auf dem Kolloquium Doktorandentag at Otto von Guericke University Magdeburg, Faculty of Computer Science, Magdeburg, Germany.
Curriculum Vitae
Since 03/2021

Acting Head of Department 4