<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T18:15:33Z</responseDate><request verb="GetRecord" identifier="oai:researchcommons.waikato.ac.nz:10289/13843" metadataPrefix="uketd_dc">https://researchcommons.waikato.ac.nz/server/oai/request</request><GetRecord><record><header><identifier>oai:researchcommons.waikato.ac.nz:10289/13843</identifier><datestamp>2020-11-12T01:08:37Z</datestamp><setSpec>com_10289_2222</setSpec><setSpec>col_10289_2223</setSpec></header><metadata><uketd_dc:uketddc xmlns:uketd_dc="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xmlns:uketdterms="http://naca.central.cranfield.ac.uk/ethos-oai/terms/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://naca.central.cranfield.ac.uk/ethos-oai/2.0/ http://naca.central.cranfield.ac.uk/ethos-oai/2.0/uketd_dc.xsd">
   <dc:title>Autoencoder-based techniques for improved classification in settings with high dimensional and small sized data</dc:title>
   <dc:creator>Daoud, Maisa</dc:creator>
   <uketdterms:advisor>Mayo, Michael</uketdterms:advisor>
   <uketdterms:advisor>Cunnigham, Sally Jo</uketdterms:advisor>
   <uketdterms:advisor>Smith, Tony C.</uketdterms:advisor>
   <dcterms:abstract>Neural network models have been widely tested and analysed usinglarge sized high dimensional datasets.  In real world application prob-lems,  the available datasets are often limited in size due to reasonsrelated to the cost or difficulties encountered while collecting the data.This  limitation  in  the  number  of  examples  may  challenge  the  clas-sification  algorithms  and  degrade  their  performance.   A  motivatingexample for this kind of problem is predicting the health status of atissue given its gene expression, when the number of samples availableto learn from is very small.Gene expression data has distinguishing characteristics attracting themachine  learning  research  community.   The  high  dimensionality  ofthe data is one of the integral features that has to be considered whenbuilding predicting models.  A single sample of the data is expressedby thousands of gene expressions compared to the benchmark imagesand texts that only have a few hundreds of features and commonlyused for analysing the existing models.  Gene expression data samplesare  also  distributed  unequally  among  the  classes;  in  addition,  theyinclude noisy features which degrade the prediction accuracy of themodels.   These  characteristics  give  rise  to  the  need  for  using  effec-tive dimensionality reduction methods that are able to discover thecomplex relationships between the features such as the autoencoders.&#xd;
This  thesis  investigates  the  problem  of  predicting  from  small  sizedhigh  dimensional  datasets  by  introducing  novel  autoencoder-basedtechniques  to  increase  the  classification  accuracy  of  the  data.   Twoautoencoder-based  methods  for  generating  synthetic  data  examplesand synthetic representations of the data were respectively introducedin the first stage of the study.  Both of these methods are applicableto the testing phase of the autoencoder and showed successful in in-creasing the predictability of the data.Enhancing the autoencoder’s ability in learning from small sized im-balanced  data  was  investigated  in  the  second  stage  of  the  projectto come up with techniques that improved the autoencoder’s gener-ated  representations.   Employing  the  radial  basis  activation  mecha-nism used in radial-basis function networks, which learn in a super-vised manner, was a solution provided by this thesis to enhance therepresentations learned by unsupervised algorithms.  This techniquewas later applied to stochastic variational autoencoders and showedpromising results in learning discriminating representations from thegene expression data.The contributions of this thesis can be described by a number of differ-ent methods applicable to different stages (training and testing) anddifferent autoencoder models (deterministic and stochastic) which, in-dividually, allow for enhancing the predictability of small sized highdimensional datasets compared to well known baseline methods.</dcterms:abstract>
   <uketdterms:institution>The University of Waikato</uketdterms:institution>
   <dcterms:issued>2020</dcterms:issued>
   <dc:type>Thesis</dc:type>
   <dc:language xsi:type="dcterms:ISO639-2">en</dc:language>
   <dcterms:isReferencedBy>https://hdl.handle.net/10289/13843</dcterms:isReferencedBy>
   <dc:identifier xsi:type="dcterms:URI">https://researchcommons.waikato.ac.nz/bitstreams/866d4b4f-d184-498e-a079-96282a8b7249/download</dc:identifier>
   <uketdterms:checksum xsi:type="uketdterms:MD5">35e058a465f39ea34a5009e60682d562</uketdterms:checksum>
   <dcterms:license>https://researchcommons.waikato.ac.nz/bitstreams/b88a8e0a-103c-4992-a8bf-30370b13bec5/download</dcterms:license>
   <uketdterms:checksum xsi:type="uketdterms:MD5">e14202ab27e47ddb00d33097327ba050</uketdterms:checksum>
   <dcterms:hasFormat>https://researchcommons.waikato.ac.nz/bitstreams/6f1d0eb0-56a4-4067-8fd8-41985977e7ac/download</dcterms:hasFormat>
   <uketdterms:checksum xsi:type="uketdterms:MD5">c828ca907fb4ab851ee8e1c8395ed90f</uketdterms:checksum>
   <dc:rights>All items in Research Commons are provided for private study and research purposes and are protected by copyright with all rights reserved unless otherwise indicated.</dc:rights>
   <dc:subject>Autoencoder, cancer prediction models</dc:subject>
   <dc:subject>thesis with publication</dc:subject>
</uketd_dc:uketddc></metadata></record></GetRecord></OAI-PMH>