STT-LLM is a structural-temporal tokenization framework that adapts LLMs to longitudinal clinical analysis without modifying their backbone architectures. It constructs biologically grounded structural-temporal embeddings and transforms them into LLM-compatible tokens through a specialised token-evolution mechanism. Evaluated on real-world longitudinal athlete datasets, it consistently improves over native LLM tokenization for sequence prediction and anomaly detection and provides contextual reasoning that aligns closely with expert assessments.
ICML 2026
AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design
Rahman, S., Rahman, M.R.
In Workshop on Generative and Agentic AI for Biology: International Conference on Machine Learning (ICML 2026)
AgentPLM uses reasoning-augmented decoding to design protein sequences with agentic protein language models.
The substitution of a urine sample that may result in an adverse analytical finding with a previously collected clean sample is strictly prohibited under WADA regulations and is referred to as sample swapping. We propose a similarity-detection framework based on a convolutional network that explicitly accounts for pattern complexity in urinary steroid profiles, evaluated on 67,651 steroid profiles collected between 2021 and 2023 on both synthetic and laboratory-confirmed similar samples.
DiGAN integrates latent diffusion modelling with an attention-guided convolutional network. The diffusion model synthesises realistic longitudinal neuroimaging trajectories from limited training data, enriching temporal context and improving robustness to unevenly spaced visits, while the attention-convolutional layer captures discriminative structural-temporal patterns that distinguish cognitively normal subjects from those with mild cognitive impairment and subjective cognitive decline. Experiments on ADNI show that DiGAN outperforms state-of-the-art baselines.
2025
PhD Thesis
Anomaly Detection in Longitudinal Clinical Profile
Doctoral thesis on anomaly detection in longitudinal clinical profiles: domain-knowledge integration, structural-temporal modelling (SACNN, STT-LLM) and interpretable reasoning for the analysis of biological samples.
We introduce Metabolism Pathway-driven Prompting (MPP), which integrates metabolic pathway information into LLM prompts to better capture structural and temporal changes in biological samples, and apply it to doping detection in sports using real-world athlete steroid data.
Sample swapping is a potential practice performed by athletes to swap doped samples with clean samples to evade positive doping tests. SACNN is a self attention-based convolutional neural network that incorporates both spatial and temporal behaviour of the longitudinal profile and generates embedding maps for fraud detection in sports, outperforming state-of-the-art baselines for sequential anomaly detection.
Generative modelling (GANs) is used to synthesise realistic blood-sample data that augments scarce anti-doping datasets and improves downstream detection models.
WITS 2024
RAG for Effective Supply Chain Security Questionnaire Automation
Reza, Z.B., Syed, A.R., Iqbal, O., Mensah, E., Liu, Q., Rahman, M.R., Maass, W.
In Proceedings of Workshop on Information Technology and Systems (WITS 2024).
WITS 2024
Towards Objectively Interpretable Fault Diagnosis for Time-Series Data in Grinding
Chan, T.T., Lange, K., Liu, R., Wein, A., Keßler, N., Rahman, M.R., Maass, W.
In Proceedings of Workshop on Information Technology and Systems (WITS 2024).
2023
AAAI 2023
SNOOP Method: Faithfulness of Text Summarizations for Single Nucleotide Polymorphisms
Maass, W., Agnes, C.K., Rahman, M.R., Almeida, J.S.
In Proceedings of the Association for the Advancement of Artificial Intelligence (AAAI) Summer Symposium 2023.
We present a data-analytical methodology supporting anti-doping decision-makers on athlete disambiguation tasks. The model helps identify swapped samples and outperforms current state-of-the-art methods and baseline models on real-world sample swapping cases.
A comparison of machine-learning algorithms combined with statistical analysis to identify erythropoietin in blood samples at sea level and moderate altitude; ensemble methods such as random forest and XGBoost provide effective tools for anti-doping organisations.