Data Stream Classification Guided by Clustering on Nonstationary Environments and Extreme Verification Latency
Data Stream Classification Guided by Clustering on Nonstationary Environments and Extreme Verification Latency
Date
2015
Authors
Souza,VMAd
Silva,DF
João Gama
Batista,GEAPA
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Data stream classification algorithms for nonstationary environments frequently assume the availability of class labels, instantly or with some lag after the classification. However, certain applications, mainly those related to sensors and robotics, involve high costs to obtain new labels during the classification phase. Such a scenario in which the actual labels of processed data are never available is called extreme verification latency. Extreme verification latency requires new classification methods capable of adapting to possible changes over time without external supervision. This paper presents a fast, simple, intuitive and accurate algorithm to classify nonstationary data streams in an extreme verification latency scenario, namely Stream Classification Algorithm Guided by Clustering - SCARGC. Our method consists of a clustering followed by a classification step applied repeatedly in a closed loop fashion. We show in several classification tasks evaluated in synthetic and real data that our method is faster and more accurate than the state-of-the-art. Copyright © SIAM.