research
2026
- KDD
Explainable AI for cancer drug response prediction: beyond univariate feature attributionsMartino Ciaperoni, Margherita Lalli, Simone Piaggesi, and 6 more authorsIn Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, 2026Predicting cancer drug response from transcriptomic profiles is a cornerstone of precision oncology, yet the scientific value of machine learning models hinges not solely on predictive accuracy, but also on their capacity to generate reliable biological insights. Current explainability approaches in this setting are computationally costly, lack robustness, and reduce complex drug response to univariate gene importance scores, overlooking the coordinated gene activity that drives sensitivity and resistance. In this work, we present ILLUME+, a scalable post-hoc explainability framework that moves beyond single-gene assessments to capture multiple, complementary forms of explanation. Integrated into our end-to-end pipeline, ILLUME+ produces more stable gene importance scores than existing baselines, recovers established drug-gene associations and mechanisms of action, and enables AI-assisted hypothesis generation to uncover novel interaction-driven molecular signals in cancer biology.
@inproceedings{ciaperoni2026explainable, title = {Explainable {AI} for cancer drug response prediction: beyond univariate feature attributions}, author = {Ciaperoni, Martino and Lalli, Margherita and Piaggesi, Simone and Varisco, Martina and Carli, Francesco and Guidotti, Riccardo and Pedreschi, Dino and Raimondi, Francesco and Giannotti, Fosca}, booktitle = {Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2}, publisher = {ACM}, year = {2026}, pages = {10749--10760}, doi = {10.1145/3770855.3819010}, } - In preparationLearning instance-specific neural decision trees with hypernetworksSimone Piaggesi, Martina Cinquini, Francesco Spinnato, and 1 more author2026In preparation
- In preparationSelf-explainable predictive models for tabular data with interpretable meta-encodingSimone Piaggesi, Riccardo Guidotti, Fosca Giannotti, and 1 more author2026In preparation
- In preparationA survey and benchmarking of differentiable treesRiccardo Guidotti, Martina Cinquini, Simone Piaggesi, and 1 more author2026In preparation
2025
- ICDM
Explanations go linear: post-hoc explainability for tabular data with interpretable meta-encodingSimone Piaggesi, Riccardo Guidotti, Fosca Giannotti, and 1 more authorIn 2025 IEEE International Conference on Data Mining (ICDM), 2025Post-hoc explainability is essential for understanding black-box machine learning models. Surrogate-based techniques are widely used for local and global model-agnostic explanations but have significant limitations. Local surrogates capture non-linearities but are computationally expensive and sensitive to parameters, while global surrogates are more efficient but struggle with complex local behaviors. In this paper, we present ILLUME, a flexible and interpretable framework grounded in representation learning, that can be integrated with various surrogate models to provide explanations for any black-box classifier. Specifically, our approach combines a globally trained surrogate with instance-specific linear transformations learned with a meta-encoder to generate both local and global explanations. Through extensive empirical evaluations, we demonstrate the effectiveness of ILLUME in producing feature attributions and decision rules that are not only accurate but also robust and computationally efficient, thus providing a unified explanation framework that effectively addresses the limitations of traditional surrogate methods.
@inproceedings{piaggesi2025linear, title = {Explanations go linear: post-hoc explainability for tabular data with interpretable meta-encoding}, author = {Piaggesi, Simone and Guidotti, Riccardo and Giannotti, Fosca and Pedreschi, Dino}, booktitle = {2025 IEEE International Conference on Data Mining (ICDM)}, publisher = {IEEE}, year = {2025}, pages = {663--672}, doi = {10.1109/ICDM65498.2025.00074}, } - TMLR
Disentangled and self-explainable node representation learningSimone Piaggesi, André Panisson, and Megha KhoslaTransactions on Machine Learning Research, 2025Node embeddings are low-dimensional vectors that capture node properties, typically learned through unsupervised structural similarity objectives or supervised tasks. While recent efforts have focused on post-hoc explanations for graph models, intrinsic interpretability in unsupervised node embeddings remains largely underexplored. To bridge this gap, we introduce DiSeNE (Disentangled and Self-Explainable Node Embedding), a framework that learns self-explainable node representations in an unsupervised fashion. By leveraging disentangled representation learning, DiSeNE ensures that each embedding dimension corresponds to a distinct topological substructure of the graph, thus offering clear, dimension-wise interpretability. We introduce new objective functions grounded in principled desiderata, jointly optimizing for structural fidelity, disentanglement, and human interpretability. Additionally, we propose several new metrics to evaluate representation quality and human interpretability. Extensive experiments on multiple benchmark datasets demonstrate that DiSeNE not only preserves the underlying graph structure but also provides transparent, human-understandable explanations for each embedding dimension.
@article{piaggesi2025disentangled, title = {Disentangled and self-explainable node representation learning}, author = {Piaggesi, Simone and Panisson, Andr{\'e} and Khosla, Megha}, journal = {Transactions on Machine Learning Research}, issn = {2835-8856}, year = {2025}, }
2024
- TKDE
DINE: dimensional interpretability of node embeddingsSimone Piaggesi, Megha Khosla, André Panisson, and 1 more authorIEEE Transactions on Knowledge and Data Engineering, 2024Graphs are ubiquitous due to their flexibility in representing social and technological systems as networks of interacting elements. Graph representation learning methods, such as node embeddings, are powerful approaches to map nodes into a latent vector space, allowing their use for various graph tasks. Despite their success, only few studies have focused on explaining node embeddings locally. Moreover, global explanations of node embeddings remain unexplored, limiting interpretability and debugging potentials. We address this gap by developing human-understandable explanations for dimensions in node embeddings. Towards that, we first develop new metrics that measure the global interpretability of embedding vectors based on the marginal contribution of the embedding dimensions to predicting graph structure. We say that an embedding dimension is more interpretable if it can faithfully map to an understandable sub-structure in the input graph - like community structure. Having observed that standard node embeddings have low interpretability, we then introduce DINE (Dimension-based Interpretable Node Embedding), a novel approach that can retrofit existing node embeddings by making them more interpretable without sacrificing their task performance. We conduct extensive experiments on synthetic and real-world graphs and show that we can simultaneously learn highly interpretable node embeddings with effective performance in link prediction.
@article{piaggesi2024dine, title = {{DINE}: dimensional interpretability of node embeddings}, author = {Piaggesi, Simone and Khosla, Megha and Panisson, Andr{\'e} and Anand, Avishek}, journal = {IEEE Transactions on Knowledge and Data Engineering}, publisher = {Institute of Electrical and Electronics Engineers (IEEE)}, volume = {36}, number = {12}, year = {2024}, pages = {7986--7997}, doi = {10.1109/TKDE.2024.3425460}, } - IEEE Access
Counterfactual and prototypical explanations for tabular data via interpretable latent spaceSimone Piaggesi, Francesco Bodria, Riccardo Guidotti, and 2 more authorsIEEE Access, 2024Artificial Intelligence decision-making systems have dramatically increased their predictive power in recent years, beating humans in many different specific tasks. However, with increased performance has come an increase in the complexity of the black-box models adopted by the AI systems, making them entirely obscure for the decision process adopted. Explainable AI is a field that seeks to make AI decisions more transparent by producing explanations. In this paper, we propose CP-ILS, a comprehensive interpretable feature reduction method for tabular data capable of generating Counterfactual and Prototypical post-hoc explanations using an Interpretable Latent Space. CP-ILS optimizes a transparent feature space whose similarity and linearity properties enable the easy extraction of local and global explanations for any pre-trained black-box model, in the form of counterfactual/prototype pairs. We evaluated the effectiveness of the created latent space by showing its capability to preserve pair-wise similarities like well-known dimensionality reduction techniques. Moreover, we assessed the quality of counterfactuals and prototypes generated with CP-ILS against state-of-the-art explainers, demonstrating that our approach obtains more robust, plausible, and accurate explanations than its competitors under most experimental conditions.
@article{piaggesi2024counterfactual, title = {Counterfactual and prototypical explanations for tabular data via interpretable latent space}, author = {Piaggesi, Simone and Bodria, Francesco and Guidotti, Riccardo and Giannotti, Fosca and Pedreschi, Dino}, journal = {IEEE Access}, publisher = {Institute of Electrical and Electronics Engineers (IEEE)}, volume = {12}, year = {2024}, pages = {168983--169000}, doi = {10.1109/ACCESS.2024.3496114}, }
2023
- PhD ThesisLearning representations for graph-structured socio-technical systemsSimone PiaggesiUniversity of Bologna, 2023
The recent widespread use of social media platforms and web services has led to a vast amount of behavioral data that can be used to model socio-technical systems. A significant part of this data can be represented as graphs or networks, which have become the prevalent mathematical framework for studying the structure and the dynamics of complex interacting systems. However, analyzing and understanding these data presents new challenges due to their increasing complexity and diversity. For instance, the characterization of real-world networks includes the need of accounting for their temporal dimension, together with incorporating higher-order interactions beyond the traditional pairwise formalism. The ongoing growth of AI has led to the integration of traditional graph mining techniques with representation learning and low-dimensional embeddings of networks to address current challenges. These methods capture the underlying similarities and geometry of graph-shaped data, generating latent representations that enable the resolution of various tasks, such as link prediction, node classification, and graph clustering. As these techniques gain popularity, there is even a growing concern about their responsible use. In particular, there has been an increased emphasis on addressing the limitations of interpretability in graph representation learning. This thesis contributes to the advancement of knowledge in the field of graph representation learning and has potential applications in a wide range of complex systems domains. We initially focus on forecasting problems related to face-to-face contact networks with time-varying graph embeddings. Then, we study hyperedge prediction and reconstruction with simplicial complex embeddings. Finally, we analyze the problem of interpreting latent dimensions in node embeddings for graphs. The proposed models are extensively evaluated in multiple experimental settings and the results demonstrate their effectiveness and reliability, achieving state-of-the-art performances and providing valuable insights into the properties of the learned representations.
@phdthesis{piaggesi2023thesis, title = {Learning representations for graph-structured socio-technical systems}, author = {Piaggesi, Simone}, school = {University of Bologna}, year = {2023}, doi = {10.48676/UNIBO/AMSDOTTORATO/10891}, }
2022
- LoG
Effective higher-order link prediction and reconstruction from simplicial complex embeddingsSimone Piaggesi, André Panisson, and Giovanni PetriIn The First Learning on Graphs Conference, 2022Methods that learn graph topological representations are becoming the usual choice to extract features to help solve machine learning tasks on graphs. In particular, low-dimensional encoding of graph nodes can be exploited in tasks such as link prediction and network reconstruction, where pairwise node embedding similarity is interpreted as the likelihood of an edge incidence. The presence of polyadic interactions in many real-world complex systems is leading to the emergence of representation learning techniques able to describe systems that include such polyadic relations. Despite this, their application on estimating the likelihood of tuple-wise edges is still underexplored. Here we focus on the reconstruction and prediction of simplices (higher-order links) in the form of classification tasks, where the likelihood of interacting groups is computed from the embedding features of a simplicial complex. Using similarity scores based on geometric properties of the learned metric space, we show how the resulting node-level and group-level feature embeddings are beneficial to predict unseen simplices, as well as to reconstruct the topology of the original simplicial structure, even when training data contain only records of lower-order simplices.
@inproceedings{piaggesi2022effective, title = {Effective higher-order link prediction and reconstruction from simplicial complex embeddings}, author = {Piaggesi, Simone and Panisson, Andr{\'e} and Petri, Giovanni}, booktitle = {The First Learning on Graphs Conference}, year = {2022}, } - EPJ Data Sci.
Time-varying graph representation learning via higher-order skip-gram with negative samplingSimone Piaggesi and André PanissonEPJ Data Science, 2022Representation learning models for graphs are a successful family of techniques that project nodes into feature spaces that can be exploited by other machine learning algorithms. Since many real-world networks are inherently dynamic, with interactions among nodes changing over time, these techniques can be defined both for static and for time-varying graphs. Here, we show how the skip-gram embedding approach can be generalized to perform implicit tensor factorization on different tensor representations of time-varying graphs. We show that higher-order skip-gram with negative sampling (HOSGNS) is able to disentangle the role of nodes and time, with a small fraction of the number of parameters needed by other approaches. We empirically evaluate our approach using time-resolved face-to-face proximity data, showing that the learned representations outperform state-of-the-art methods when used to solve downstream tasks such as network reconstruction. Good performance on predicting the outcome of dynamical processes such as disease spreading shows the potential of this method to estimate contagion risk, providing early risk awareness based on contact tracing data.
@article{piaggesi2022timevarying, title = {Time-varying graph representation learning via higher-order skip-gram with negative sampling}, author = {Piaggesi, Simone and Panisson, Andr{\'e}}, journal = {EPJ Data Science}, publisher = {Springer Science and Business Media LLC}, volume = {11}, number = {1}, year = {2022}, doi = {10.1140/epjds/s13688-022-00344-8}, } - Front. Big Data
Mapping urban socioeconomic inequalities in developing countries through Facebook advertising dataSimone Piaggesi, Serena Giurgola, Márton Karsai, and 3 more authorsFrontiers in Big Data, 2022Ending poverty in all its forms everywhere is the number one Sustainable Development Goal of the UN 2030 Agenda. To monitor the progress toward such an ambitious target, reliable, up-to-date and fine-grained measurements of socioeconomic indicators are necessary. When it comes to socioeconomic development, novel digital traces can provide a complementary data source to overcome the limits of traditional data collection methods, which are often not regularly updated and lack adequate spatial resolution. In this study, we collect publicly available and anonymous advertising audience estimates from Facebook to predict socioeconomic conditions of urban residents, at a fine spatial granularity, in four large urban areas: Atlanta (USA), Bogotá (Colombia), Santiago (Chile), and Casablanca (Morocco). We find that behavioral attributes inferred from the Facebook marketing platform can accurately map the socioeconomic status of residential areas within cities, and that predictive performance is comparable in both high and low-resource settings. Our work provides additional evidence of the value of social advertising media data to measure human development and it also shows the limitations in generalizing the use of these data to make predictions across countries.
@article{piaggesi2022mapping, title = {Mapping urban socioeconomic inequalities in developing countries through Facebook advertising data}, author = {Piaggesi, Simone and Giurgola, Serena and Karsai, M{\'a}rton and Mejova, Yelena and Panisson, Andr{\'e} and Tizzoni, Michele}, journal = {Frontiers in Big Data}, publisher = {Frontiers Media SA}, volume = {5}, year = {2022}, doi = {10.3389/FDATA.2022.1006352}, }
2020
- HSScomms
Gender gaps in urban mobilityLaetitia Gauvin, Michele Tizzoni, Simone Piaggesi, and 5 more authorsHumanities and Social Sciences Communications, 2020Mobile phone data have been extensively used to study urban mobility. However, studies based on gender-disaggregated large-scale data are still lacking, limiting our understanding of gendered aspects of urban mobility and our ability to design policies for gender equality. Here we study urban mobility from a gendered perspective, combining commercial and open datasets for the city of Santiago, Chile. We analyze call detail records for a large cohort of anonymized mobile phone users and reveal a gender gap in mobility: women visit fewer unique locations than men, and distribute their time less equally among such locations. Mapping this mobility gap over administrative divisions, we observe that a wider gap is associated with lower income and lack of public and private transportation options. Our results uncover a complex interplay between gendered mobility patterns, socio-economic factors and urban affordances, calling for further research and providing insights for policymakers and urban planners.
@article{gauvin2020gender, title = {Gender gaps in urban mobility}, author = {Gauvin, Laetitia and Tizzoni, Michele and Piaggesi, Simone and Young, Andrew and Adler, Natalia and Verhulst, Stefaan and Ferres, Leo and Cattuto, Ciro}, journal = {Humanities and Social Sciences Communications}, publisher = {Springer Science and Business Media LLC}, volume = {7}, number = {1}, year = {2020}, doi = {10.1057/s41599-020-0500-x}, } - SEBD
Explainability methods for natural language processing: applications to sentiment analysisFrancesco Bodria, André Panisson, Alan Perotti, and 1 more authorIn CEUR Workshop Proceedings, 2020Sentiment analysis is the process of classifying natural language sentences as expressing positive or negative sentiments, and it is a crucial task where the explanation of a prediction might arguably be as necessary as the prediction itself. We analysed different explanation techniques, and we applied them to the classification task of Sentiment Analysis. We explored how attention-based techniques can be exploited to extract meaningful sentiment scores with a lower computational cost than existing XAI methods.
@inproceedings{bodria2020explainability, title = {Explainability methods for natural language processing: applications to sentiment analysis}, author = {Bodria, Francesco and Panisson, Andr{\'e} and Perotti, Alan and Piaggesi, Simone}, booktitle = {CEUR Workshop Proceedings}, volume = {2646}, pages = {100--107}, year = {2020}, organization = {CEUR-WS}, }
2019
- CVPRW
Predicting city poverty using satellite imagerySimone Piaggesi, Laetitia Gauvin, Michele Tizzoni, and 7 more authorsIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019Reliable data about socio-economic conditions of individuals, such as health indexes, consumption expenditures and wealth assets, remain scarce for most countries. Traditional methods to collect such data include on site surveys that can be expensive and labour intensive. On the other hand, remote sensing data, such as high-resolution satellite imagery, are becoming largely available. To circumvent the lack of socio-economic data at high granularity, computer vision has already been applied successfully to raw satellite imagery sampled from resource poor countries. In this work we apply a similar approach to the metropolitan areas of five different cities in North and South America, starting from pre-trained convolutional models used for poverty mapping in developing regions. Applying a transfer learning process we estimate household income from visual satellite features. The urban environment we consider is characterized by different features with respect to the resource-poor training environment, such as the high heterogeneity in population density. By leveraging both official and crowd-sourced data at city scale, we show the feasibility of estimating the socio-economic conditions of different neighborhoods from satellite data.
@inproceedings{piaggesi2019predicting, title = {Predicting city poverty using satellite imagery}, author = {Piaggesi, Simone and Gauvin, Laetitia and Tizzoni, Michele and Cattuto, Ciro and Adler, Natalia and Verhulst, Stefaan and Young, Andrew and Price, Rhiannan and Ferres, Leo and Panisson, Andr{\'e}}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops}, year = {2019}, }