Flakiness, or non-deterministic behavior of builds and tests, is a major and continuing impediment to software testing. The Flaky Tests Workshop (FTW) aims to bring together industry professionals and academics to discuss the wide effects of test flakiness, propose solution approaches, report on the industrial state of practice, and present empirical studies, as well as case studies.
Status
The International Flaky Tests Workshop (FTW) will continue in 2027. Following FTW 2024, FTW 2025, and FTW 2026, we invite submissions on test flakiness and non-determinism in testing.
Extended abstract (max. 2 pages including references): New ideas, problems and challenges, viewpoints, work in progress. Extended abstracts are free of APC (article processing charge).
Short paper (max. 5 pages including references): Technical research, experience reports, empirical studies.
Organizing Committee

Sebastian Baltes is a Professor of Software Engineering at Heidelberg University, Germany, and an Adjunct Professor at the University of Adelaide, Australia. He received his Ph.D. in Computer Science from the University of Trier, Germany, in 2019 and spent 3.5 years in the software industry before returning to academia. His research focuses on software analytics, that is, processing, analyzing, and visualizing software engineering data to monitor, govern, and improve development processes and tools. He is also interested in interdisciplinary research and the methodological foundations of empirical software engineering. For him, thoroughly understanding the state of practice is an essential first step toward improving how software is developed. More broadly, he believes software engineering must become more evidence-based, which is only possible if academic researchers provide actionable insights on topics relevant to practitioners. His work has been published in leading venues including ICSE, FSE, TSE, TOSEM, and EMSE, has received multiple best paper awards, including two ACM SIGSOFT Distinguished Paper Awards, and has received funding from Google and SAP. More information is available on his web page.

Fabian Leinen is a PhD researcher at the Technical University of Munich, focused on regression tests. He specializes in analyzing and managing flaky tests in continuous integration environments. He aims to enhance the understanding and practical handling of test flakiness in software development.

Joanna Kisaakye is a PhD researcher at the University of Antwerp. She is part of the Lab on Reengineering (LoRE). Her research centres around software evolution and software quality. Currently, she is focusing on flakiness scores, their adoption and adaptations for practice.

Shanto Rahman is an incoming Assistant Professor in the Department of Computer Science at The University of Texas at Dallas. She is currently completing her PhD in Electrical and Computer Engineering at The University of Texas at Austin, advised by Prof. August Shi. Her research focuses on software testing, with a particular emphasis on flaky tests. Over the years, she has developed techniques for predicting, debugging, and repairing flaky tests.
Steering Committee

August Shi is an Assistant Professor in the Chandra Family Department of Electrical and Computer Engineering at the University of Texas at Austin. August has worked extensively in the area of flaky test research, publishing work on detecting, debugging, and repairing flaky tests. For his work on flaky tests, August has been awarded a SIGSOFT Outstanding Doctoral Dissertation Award and an NSF CAREER Award. August received his PhD in Computer Science from the University of Illinois at Urbana-Champaign.

Wing Lam is an Assistant Professor in the Computer Science department at George Mason University. Dr. Lam works on several topics in software engineering, with a focus on software testing. His research improves software dependability by characterizing bugs and developing novel techniques to detect and tame bugs. Wing has published in top-tier conferences, such as ESEC/FSE, ICSE, ISSTA, OOPSLA, and TACAS. His techniques have helped detect and fix bugs in open-source projects and have impacted how Dragon Testing, Microsoft, and Tencent developers test their code. Wing has been the recipient of several awards, including an ACM SIGSOFT Outstanding Doctoral Dissertation Award and an NSF CAREER award. More information is available on his web page.

Owain Parry is an early-career researcher working on flaky tests. He published multiple papers in this field, including an extensive systematic literature review on the topic, a developer survey, and a machine learning-driven technique for detecting flaky tests.

Martin Gruber is a Software Engineer working with the BMW Group, where he is concerned with software quality and testing in CI environments. His dissertation focused on understanding and mitigating test flakiness.

Tim A. D. Henderson is Staff Software Engineer at Google working on Google’s proprietary CI platform: the Test Automation Platform (TAP). Dr. Henderson specializes in testing, fault localization, semantic code clone detection, program analysis, graph mining, and databases.

Gordon Fraser is a full professor in Computer Science at the University of Passau, Germany. He received a PhD in computer science from Graz University of Technology, Austria, in 2007, worked as a post-doc at Saarland University, and was a Senior Lecturer at the University of Sheffield, UK. The central theme of his research is improving software quality, and his recent research concerns the prevention, detection, and removal of defects in software. He has been facing flaky tests in particular since the creation of the automated test generation tool EvoSuite. He has chaired major software engineering conferences (e.g., ASE, ISSTA, ICST) and workshops (e.g., Mutation, A-MOST, AST, TAIC-PART, CSI-SE, CSTVA, Gamify).

Phil McMinn has been a Lecturer in the Computer Science department at the University of Sheffield since 2006. He was awarded his PhD in 2005, which was funded by DaimlerChrysler Research and Technology. He has published several papers in the field of search-based testing. His research interests cover software engineering, with a particular focus on software testing, program transformation and agent-based systems and modelling. His research has been funded by the UK Engineering and Physical Science Research Council (EPSRC) to work on reducing oracles costs of testing, testing techniques for agent-based systems and the automatic reverse engineering of state machine descriptions from software.