Computing and Mathematical Sciences Papers
Permanent URI for this collectionhttps://researchcommons.waikato.ac.nz/handle/10289/6
This collection houses research from the School of Computing and Mathematical Sciences at the University of Waikato.
Browse
Recent Submissions
Item type: Item , Periodic boundary conditions and G2 cosmology(IOP Publishing, 2023) Coley, Alan A.; Lim, Woei ChetIn the standard concordance cosmology the spatial curvature is assumed to be constant and zero (or at least very small). In particular, in numerical computations of the structure of the universe using N-body simulations, exact periodic boundary conditions are assumed which constrains the spatial curvature. In order to confirm this qualitatively, we numerically evolve a special class of spatially inhomogeneous G2 models with both periodic initial data and non periodic initial data using zooming techniques. We consequently demonstrate that in these models periodic initial conditions do indeed suppress the growth of the spatial curvature as the models evolve away from their initial isotropic and spatially homogeneous state, thereby verifying that the spatial curvature is necessarily very small in standard cosmology.Item type: Item , Automatic detection of Android crypto ransomware using supervisor reduction(Springer Nature, 2024) Chew, Christopher J.W.; Malik, Robi; Kumar, Vimal; Patros, PanosThis paper proposes a finite-state machine based approach to recognise crypto ransomware based on their behaviour. Malicious and benign Android applications are executed to capture the system calls they generate, which are then filtered and tokenised and converted to finite-state machines. The finite-state machines are simplified using supervisor reduction, which generalises the behavioural patterns and produces compact classification models. The classification models can be implemented in a lightweight monitoring system to detect malicious behaviour of running applications quickly. An extensive set of cross validation experiments is carried out to demonstrate the viability of the approach, which show that ransomware can be classified accurately with an F1 score of up to 93.8%.Item type: Item , Enriching cultural heritage communities: New tools and technologies(Oxford University Press, 2024) Dix, Alan; Jones, Elizabeth; Cowgill, Rachel; Armstrong, Charlotte; Ridgewell, Rupert; Twidale, Michael B.; Downie, J. Stephen; Reagan, Maureen; Bashford, Christina; Bainbridge, David; Neads, Carys-Ann; Davies, VinceThis paper explores ways in which scholarly skill and expertise might be embodied in tools and sustainable practices that enable communities to create and manage their own digital archives. We focus particularly on tools and practices related to the recording and annotation of digitized materials. The paper is based on co-production practice in two very different kinds of community. Although the communities are different we find that tools designed for a specific community are valuable for others, thus offering the promise of general tools to support community-centred digitization and potentially also traditional archival practice.Item type: Item , A multimodal bracelet to acquire muscular activity and gyroscopic data to study sensor fusion for intent detection(MDPI, 2024) Andreas, Daniel; Hou, Zhongshi; Tabak, Mohamad Obada; Dwivedi, Anany; Beckerle, PhilippResearchers have attempted to control robotic hands and prostheses through biosignals but could not match the human hand. Surface electromyography records electrical muscle activity using non-invasive electrodes and has been the primary method in most studies. While surface electromyography-based hand motion decoding shows promise, it has not yet met the requirements for reliable use. Combining different sensing modalities has been shown to improve hand gesture classification accuracy. This work introduces a multimodal bracelet that integrates a 24-channel force myography system with six commercial surface electromyography sensors, each containing a six-axis inertial measurement unit. The device’s functionality was tested by acquiring muscular activity with the proposed device from five participants performing five different gestures in a random order. A random forest model was then used to classify the performed gestures from the acquired signal. The results confirmed the device’s functionality, making it suitable to study sensor fusion for intent detection in future studies. The results showed that combining all modalities yielded the highest classification accuracies across all participants, reaching (Formula presented.) on average, effectively reducing misclassifications by 37% and 22% compared to using surface electromyography and force myography individually as input signals, respectively. This demonstrates the potential benefits of sensor fusion for more robust and accurate hand gesture classification and paves the way for advanced control of robotic and prosthetic hands.Item type: Item , Topological and dynamic characteristics in forced anisotropic magnetohydrodynamic turbulence(AIP Publishing, 2026) Gao, Kai; Jiang, Bin; Li, Cheng; Yang, Yan; Zhou, Kangcheng; Matthaeus, William H.; Oughton, Sean; Wan, MinpingThe invariants of the velocity gradient tensor in turbulence offer a compact description of local kinematics and flow topology. For incompressible magnetohydrodynamic turbulence, analysis of the second and third invariants (Q, R) of the velocity gradient tensor clarifies how coherent structures are organized and evolve. Extending the same analysis to the magnetic field gradient tensor provides additional information on the dynamics. In this study, pseudo-spectral simulation is used to obtain the velocity and magnetic field of the turbulent flow, and analysis of the flow field is conducted through joint probability density functions (PDFs) of the invariants. Furthermore, we explore the influence of the external mean magnetic field strength, B 0. The results show that when an external magnetic field is present, the Q–R joint PDF no longer maintains the familiar teardrop distribution for the velocity field, and the flow field structure tends to be two-dimensional with increasing B 0. For the fluctuation magnetic field, the Q–R joint PDF takes on a “cigar” shape that becomes more elongated as B 0 increases. Moreover, as the strength of the external mean magnetic field increases, the turbulence exhibits enhanced small-scale dissipation and localization, accompanied by a reduction in the effective dimensionality of the system toward a quasi-two-dimensional regime.Item type: Item , Anomaly detection for evolving maritime trajectories with continual learning(Springer Nature, 2026) Julian, Jack; Koh, Yun Sing; Bifet, AlbertAnomaly detection in live trajectory data is a critical task for ensuring safety, security, and legality in global transport. Traditional anomaly detection methods often struggle with dynamic and evolving trajectory patterns, especially as systems must adapt to new scenarios over time due to increased traffic, geopolitical events, or global warming. We propose a continual learning approach to detect anomalous activity in moving vessels. Unlike conventional static models, our method leverages continual learning to enable the model to learn from new data continuously and recognise specific behaviours dependent on position and recent movements. We implement an adapter-based framework, Continual Learning for AIS Anomalies (CLAISA), that adapts to shifting behavioural environments in transportation, ensuring the system can identify novel and evolving patterns of anomalies, such as deviations from expected routes, irregular speed changes, or unusual local movements. Evaluations on synthetic maritime trajectory datasets spanning sparsely populated waters and heavily trafficked shipping lanes demonstrate that CLAISA achieves up to a decrease in error for trajectory forecasting and consistently outperforms benchmark methods in anomaly detection on synthetically generated datasets.Item type: Item , Integrating agentic artificial intelligence into pasture-based dairy systems: Applications, governance, and future directions(Elsevier, 2026) Eastwood, Callum; Lim, Nick Jin Sean; Durie, Rachel; Dela Rue, Brian; Bifet, Albert; Reed, CharlotteDigitalization and artificial intelligence (AI) provide opportunities for improved management of pasture-based dairy systems. Agentic AI, where autonomous systems can perceive and act independently of humans, are potentially transformative. This mini-review explores the potential use of agentic AI to address key challenges in pasture-based dairy systems. Applications include autonomous grazing animal health prediction, environmental modeling, and virtual assistants. Agentic AI can integrate multiple data sources to support real-time, farm-specific decisions. It also presents opportunities for employee training, enhanced advisory services, and digital twin modeling. However, deployment of agentic AI introduces governance, ethical, and socio-technical considerations. Issues of data ownership, transparency, explainability, and trust must be addressed. The experiential and tacit knowledge of farmers must be integrated into AI systems through hybrid intelligence (human and AI). Oversight models ranging from Human-in-the-Loop to Human-in-Command are necessary to ensure safe and responsible use. Future research should focus not only on technical AI development, but farmer-centered design, robust assurance frameworks, and inclusive and responsible innovation ecosystems that align technological progress with dairy sector, civil society, values, and needs.Item type: Item , An agentic system for LLM-driven public transportation analytics: A practical application and case study in Salvador-Brazil(ACM, 2026) Borges, Lucas T.; Liu, Fei T.; Rios, Tatiane; Ferreira, Marcos V.; Carmo, Clovis; Coimbra, Danilo B.; Nery, Jorge; Souza, Matheus; Garcia, Noe O.; Bifet, Albert; Rios, RicardoPublic transportation agencies gather vast and heterogeneous datasets, yet their decisions still depend largely on manual queries and fragmented analyses. To address this gap, we introduce SUNTInsight, an agentic system that combines large language models, machine learning, and visual analytics to strengthen human decision-making. Through a seamless workflow, free-form prompts are automatically converted into SQL queries, relevant data are retrieved and processed, and an LLM, supported by visual interfaces, interprets the results to produce actionable findings. Applied to a case study in Salvador-Brazil, SUNTInsight revealed spatial heterogeneity, identified high-demand segments, and recommended targeted operational strategies, such as short turns and headway control, instead of broad fleet expansions. The implementation follows the principle of least privilege and runs generated code in isolated execution, mitigating risk while keeping humans in control.Item type: Item , Q-Cowrie: An adaptive honeypot to analyse attackers’ behaviour(Springer Nature, 2026) Var Naseri, Maryam; Welch, Ian; Haseeb, Junaid; Mansoori, MasoodHoneypots in computer security have been used as effective security solutions to lure attackers, capture their interactions with the honeypot systems and study their behaviour. Attackers interacting with honeypots may use Artificial Intelligence (AI)-based techniques to detect the presence of honeypots leading to evasion by the attackers. This paper discusses the application of Reinforcement Learning (RL) to address these issues by improving response generation in honeypots. We propose “Q-Cowrie”, a honeypot that is built upon customising a medium interaction server honeypot, that is, Cowrie, to increase the honeypot’s deception. RL capabilities have been integrated into the honeypot to support adaptive behaviour while interacting with attackers. Two experimental studies have been conducted in which Cowrie and Q-Cowrie honeypots were used, respectively. First, we deployed a Cowrie honeypot to capture cyber attacks and identify attackers’ goals and techniques. This allowed us to create a probabilistic model, that is, the Markov Decision Making Process (MDP), to understand the decision-making process of attackers in different situations. Learning from attackers’ unique patterns and applying RL techniques, Q-Cowrie was able to actively interact with attackers, making adaptive decisions.Item type: Item , On DR-semigroups satisfying the ample conditions(Springer Nature, 2026) Stokes, Tim E.A DR-semigroup S (also known as a reduced E-semiabundant or reduced E-Fountain semigroup) is here viewed as a semigroup equipped with two unary operations D, R satisfying finitely many equational laws. Examples include DRC-semigroups (hence Ehresmann semigroups), which also satisfy the congruence conditions. The ample conditions on DR-semigroups are studied here and are defined by the laws (Formula presented.) Two natural partial orders may be defined on a DR-semigroup and we show that the ample conditions hold if and only if the two orders are equal and the projections (elements of the form D(x)) commute with one-another. Restriction semigroups satisfy the ample conditions, but we give non-restriction examples using closure operators on sets, strongly order-preserving functions on a quasiordered set, and certain subsets of partial categories. Following the work of Stein, we show how to construct a certain partial algebra C(S) from any DR-semigroup, which is a category if S satisfies the congruence conditions, but is “almost" a category if the ample conditions hold. We then characterise the ample conditions in terms of the equality of two natural partial orders, and in terms of a converse of the condition on S ensuring that C(S) is a category. Our main result is an ESN-style theorem for DR-semigroups satisfying the ample conditions, based on the C(S) construction. We also obtain an embedding theorem, generalizing a result for restriction semigroups due to Lawson.Item type: Item , Building adaptive knowledge bases for evolving continual learning models(Springer Nature, 2025) Julian, Jack; Koh, Yun Sing; Bifet, AlbertContinual learning addresses catastrophic forgetting and knowledge transfer when learning from task streams. Dynamic architectures have introduced task-specific components like adapters layered over fixed pre-trained backbones. However, identifying the task of a new input remains a core challenge, leading to task-agnostic and dynamic detection methods. Existing approaches often overlook the reuse of previously learned adapters, missing opportunities for efficient forward and backwards transfer. We propose Continual Adapter-Based Learning (CABLE), a reinforcement learning framework that computes gradient similarity between new examples and past tasks. This similarity score drives a policy that assigns existing adapters when beneficial, rewarding improved performance and reducing reliance on newly initialised parameters. CABLE adopts a dynamic adapter routing strategy without assuming prior task labels. Evaluations on image classification and time series forecasting show that CABLE mitigates catastrophic forgetting and promotes efficient knowledge transfer across tasks.Item type: Item , The age of DDoScovery: An empirical comparison of industry and academic DDoS assessments(ACM, 2024) Hiesgen, Raphael; Nawrocki, Marcin; Barcellos, Marinho; Kopp, Daniel; Hohlfeld, Oliver; Chan, Echo; Dobbins, Roland; Doerr, Christian; Rossow, Christian; Thomas, Daniel R.; Jonker, Mattijs; Mok, Ricky; Luo, Xiapu; Kristoff, John; Schmidt, Thomas C.; Wählisch, Matthias; Claffy, K. C.Motivated by the impressive but diffuse scope of DDoS research and reporting, we undertake a multistakeholder (joint industry-academic) analysis to seek convergence across the best available macroscopic views of the relative trends in two dominant classes of attacks - direct-path attacks and reflection-amplification attacks. We first analyze 24 industry reports to extract trends and (in)consistencies across observations by commercial stakeholders in 2022. We then analyze ten data sets spanning industry and academic sources, across four years (2019-2023), to find and explain discrepancies based on data sources, vantage points, methods, and parameters. Our method includes a new approach: we share an aggregated list of DDoS targets with industry players who return the results of joining this list with their proprietary data sources to reveal gaps in visibility of the academic data sources. We use academic data sources to explore an industry-reported relative drop in spoofed reflection-amplification attacks in 2021-2022. Our study illustrates the value, but also the challenge, in independent validation of security-related properties of Internet infrastructure. Finally, we reflect on opportunities to facilitate greater common understanding of the DDoS landscape. We hope our results inform not only future academic and industry pursuits but also emerging policy efforts to reduce systemic Internet security vulnerabilities.Item type: Item , BusEnv: A multi-agent reinforcement learning environment and benchmark for urban public transportation(ACM, 2026) da Silva e Silva, Wesley; Rios, Ricardo A.; da Costa Fonseca, Rafael; Ponnambalam, Sabarikirishwaran; Cassé, Léa; dos Santos Ferreira, Marcos Vinícius; Bifet, Albert; Rios, Tatiane N.Reinforcement learning (RL) offers a powerful paradigm for managing complex, dynamic transportation systems where autonomous agents must adapt to uncertain and rapidly changing environments. We present BusEnv, a benchmark environment grounded in real-world data from the Salvador Urban Transportation Network, encompassing approximately 700,000 passengers, 2,000 vehicles, 400 lines, and 3,000 stops, collected between March 2024 and March 2025 at sub-minute resolution. BusEnv simulates realistic bus operations with stochastic passenger demand, route-specific travel times, and traffic-dependent variability, enabling controlled experimentation under partially observable, high-dimensional conditions. The reward function integrates multiple objectives, such as passenger service quality, operational efficiency, maintenance adherence, and sustainability, allowing the assessment of how different RL algorithms balance these competing factors. We evaluate nine baseline methods implemented in MARLlib, analyzing their convergence, robustness, and environmental impact when deployed under independent-learning conditions. Results show that PPO-based approaches achieve the highest stability and lowest energy waste, linking algorithmic robustness to sustainability performance. By combining data realism with reproducibility and extensibility, BusEnv establishes a foundation for systematic research on learning-based transport management and provides a scalable testbed for future studies on cooperative, sustainability-aware reinforcement learning.Item type: Item , Editorial: Wearables for human-robot interaction and collaboration(Frontiers, 2025-12-09) Zhang, X; Dwivedi, Anany; Leng, Y; Liarokapis, M; Lahr, GJGItem type: Item , Factors controlling the statistics of magnetic reconnection in magnetohydrodynamic turbulence(American Physical Society, 2026-01-22) Khan, Muhammad Bilal; Shay, Michael A.; Oughton, Sean; Matthaeus, William H.; Haggerty, C. C.; Adhikari, Subash ; Cassak, Paul A.; Fordin, S.; O'Donnell, Daniel; Yang, Yan; Bandyopadhyay, Riddhi; Roy, SohomWe study the statistics of dynamical quantities associated with magnetic reconnection events embedded in a sea of strong background magnetohydrodynamic turbulence using direct numerical simulations. We focus on the relationship of the reconnection properties to the statistics of global turbulent fields. We show that the distribution in turbulence of reconnection rates (determined by upstream fields) is strongly correlated with the magnitude of the global turbulent magnetic field at the correlation scale. The average reconnection rates, and associated dissipation rates, during turbulence are thus much larger than predicted by using turbulent magnetic field fluctuation amplitudes at the dissipation or kinetic scales. Magnetic reconnection may, therefore, be playing a major role in energy dissipation in astrophysical and heliospheric turbulence.Item type: Item , ARES: Anomaly Recognition Model for Edge Streams(ACM, 2026-04-20) Mungari, Simone; Bifet, Albert; Manco, Giuseppe; Pfahringer, BernhardMany real-world scenarios involving streaming information can be represented as temporal graphs, where data flows through dynamic changes in edges over time. Anomaly detection in this context has the objective of identifying unusual temporal connections within the graph structure. Detecting edge anomalies in real time is crucial for mitigating potential risks. Unlike traditional anomaly detection, this task is particularly challenging due to concept drifts, large data volumes, and the need for real-time response. To face these challenges, we introduce ARES, an unsupervised anomaly detection framework for edge streams. ARES combines Graph Neural Networks (GNNs) for feature extraction with Half-Space Trees (HST) for anomaly scoring. GNNs capture both spike and burst anomalous behaviors within streams by embedding node and edge properties in a latent space, while HST partitions this space to isolate anomalies efficiently. ARES operates in an unsupervised way without the need for prior data labeling. To further validate its detection capabilities, we additionally incorporate a simple yet effective supervised thresholding mechanism. This approach leverages statistical dispersion among anomaly scores to determine the optimal threshold using a minimal set of labeled data, ensuring adaptability across different domains. We validate ARES through extensive evaluations across several real-world cyber-attack scenarios, comparing its performance against existing methods while analyzing its space and time complexity. The code used to perform the experiments is publicly available at https://github.com/AnomalyRecognitionModelForEdgeStreams/ARES.Item type: Item , Adaptive approaches towards fully incremental prediction interval for data stream regression(Springer Nature, 2026) Sun, Yibin; Pfahringer, Bernhard; Gomes, Heitor Murilo; Bifet, AlbertPrediction intervals (PIs) are a practical tool for uncertainty quantification in regression, but comparatively little work has addressed fully incremental PI generation for data streams. In streaming settings, data arrive continuously, each instance is typically processed once, and concept drift can quickly invalidate a previously well-calibrated interval. These properties make many batch PI methods and window-based adaptations difficult to apply efficiently. This paper studies Adaptive Prediction Interval (AdaPI), an online post-calibration framework that adjusts interval width according to observed coverage. We instantiate the framework with a fully incremental variant of Mean and Variance Estimation (MVE) and investigate three adaptive scaling functions. We also adopt an evaluation perspective that jointly considers coverage accuracy and interval width. Experiments on a collection of real-world and synthetic regression streams show that AdaPI can often move coverage closer to the desired confidence level while maintaining competitive interval width; under the default 95% confidence setting and coverage-heavy CING weighting, the linear variant frequently gives the strongest empirical coverage–width trade-off among the three adaptive strategies.Item type: Item , Co-designing smart cities: A case study in the New Zealand context(ACM, 2025) Turner, Jessica; Jones, Ben E.; Everitt, Aluna; König, Jemma L.; Fredericks, Joel; Yoo, Soojeong; Tran, Tram Thi Minh; Pantidi, Nadia; Hoang, Thuong; Hoggenmueller, Marius; Caldwell, Glenda; Tag, Benjamin; Andres, Josh; Davis, Hilary; Boden, Marie; Zhu, Howe; Harman, Joel; Rahman, JessicaDesigning smart cities that align with community needs requires both innovation and active participation. However, engaging the community can be challenging due to the diverse range of stakeholders, their requirements and hesitation to participate. Co-design offers a collaborative solution, involving the community from the outset, and ensuring designs meet expectations. This research explores the feasibility of co-design approaches for smart city development in New Zealand, using a local council as a case study. We conducted an online survey with 248 participants to assess public attitudes. Survey results highlighted strong interest in addressing traffic congestion, environmental monitoring and public safety. Using these results we conducted a co-design workshop (n=13) to encourage ideation. The co-design process generated innovative concepts, evaluated based on cost, deliverability, and novelty. Findings demonstrate the potential of co-design methodologies to bridge the gap between community needs and smart city technology, offering a foundation for future urban development initiatives.Item type: Item , On class operators for the lower radical class and semisimple closure constructions(Springer Nature, 2024) McConnell, N.R.; McDougall, R.G.; Stokes, Tim E.; Thornton, L.K.We construct the lower radical class and the semisimple closure for a given class using class operators and detail some of the properties of these operators and their interplay with the operators already used in radical theory. The setting is the class of algebras introduced by Puczy lowski which ensures the results hold in groups, multi-operator groups such as rings, as well as loops and hoops.Item type: Item , Concept drift detection in delayed and partially labeled data streams: An experimental survey(Elsevier, 2026) Alencar, Brenno; Cassales, Guilherme; Gomes, Heitor Murilo; Prazeres, Cássio; Rios, Tatiane N.; Bifet, Albert; Rios, Ricardo A.Fully labeled data is the ideal scenario for supervised model training. However, labels might be scarce in many situations, specially when large data is produced as an open-ended stream. Moreover, in dynamic streams, it is assumed that the underlying process generating the data is non-stationary, changing its behavior over time. Therefore, concepts learned by the classification model are likely to change, and non-adaptive models will suffer performance degradation, a phenomenon known as concept drift. The challenge of recognizing and reacting to such changes is even more complex by facing streams with partial or delayed labels. This refers to realistic scenarios where the true label of an incoming point is not immediately available (delay) or might never be available (partial). In this work, we focus on investigating the performance of concept drift detectors in handling delayed and partially labeled data streams. We categorize the main methods from the literature, providing a taxonomy of concept drift detectors and an overview of the most influential approaches for evaluating these detectors under conditions of delayed or partial labeling. Finally, we conduct a series of experiments to analyze the performance of these detectors in scenarios with label scarcity. Concluding, the survey discusses its main limitations and offers insights into potential avenues for future research in the field.