<?xml version="1.0" encoding="UTF-8"?><?xml-model type="application/xml-dtd" href="http://jats.nlm.nih.gov/publishing/1.1d3/JATS-journalpublishing1.dtd"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1d3 20150301//EN" "http://jats.nlm.nih.gov/publishing/1.1d3/JATS-journalpublishing1.dtd">
<article xmlns:ali="http://www.niso.org/schemas/ali/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.1d3" specific-use="Marcalyc 1.2" article-type="research-article" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="redalyc">205</journal-id>
<journal-title-group>
<journal-title specific-use="original" xml:lang="es">Cuadernos de Administración</journal-title>
<abbrev-journal-title abbrev-type="publisher" xml:lang="es">Cuad Adm</abbrev-journal-title>
</journal-title-group>
<issn pub-type="ppub">0120-3592</issn>
<issn pub-type="epub">1900-7205</issn>
<publisher>
<publisher-name>Pontificia Universidad Javeriana</publisher-name>
<publisher-loc>
<country>Colombia</country>
<email>revistascientificasjaveriana@gmail.com</email>
</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="art-access-id" specific-use="redalyc">20562876001</article-id>
<article-id pub-id-type="doi">https://doi.org/10.11144.Javeriana.cao33.tlcms</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artículos</subject>
</subj-group>
</article-categories>
<title-group>
<article-title xml:lang="en">
<bold>Two-level clustering methodology for smart metering data<xref ref-type="fn" rid="fn1">*</xref>
</bold>
</article-title>
<trans-title-group>
<trans-title xml:lang="es">
<bold>Metodología de agrupación en dos niveles para una medición de datos inteligente</bold>
</trans-title>
</trans-title-group>
<trans-title-group>
<trans-title xml:lang="pt">
<bold>Metodologia de agrupação em dois níveis para uma medição de dados inteligente</bold>
</trans-title>
</trans-title-group>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0002-5154-4441</contrib-id>
<name name-style="western">
<surname>Arco García</surname>
<given-names>Leticia</given-names>
</name>
<xref ref-type="corresp" rid="corresp1"><sup>a</sup></xref>
<xref ref-type="aff" rid="aff1"/>
<email>larcogar@vub.be</email>
</contrib>
<contrib contrib-type="author" corresp="no">
<name name-style="western">
<surname>Casas Cardoso</surname>
<given-names>Gladys María</given-names>
</name>
<xref ref-type="aff" rid="aff2"/>
</contrib>
<contrib contrib-type="author" corresp="no">
<contrib-id contrib-id-type="orcid">http://orcid.org/0000-0001-6346-4564</contrib-id>
<name name-style="western">
<surname>Nowé</surname>
<given-names>Ann</given-names>
</name>
<xref ref-type="aff" rid="aff3"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution content-type="original">Vrije Universiteit Brussels, Belgium</institution>
<institution content-type="orgname">Vrije Universiteit Brussels</institution>
<country country="NL">Países bajos</country>
</aff>
<aff id="aff2">
<institution content-type="original">CENSA International College, United States</institution>
<institution content-type="orgname">CENSA International College</institution>
<country country="US">Estados Unidos</country>
</aff>
<aff id="aff3">
<institution content-type="original">Vrije Universiteit Brussels, Belgium</institution>
<institution content-type="orgname">Vrije Universiteit Brussels</institution>
<country country="NL">Países bajos</country>
</aff>
<author-notes>
<corresp id="corresp1"><sup>a</sup> Corresponding author. E-mail: <email>larcogar@vub.be</email>
</corresp>
</author-notes>
<pub-date pub-type="epub-ppub">
<season>Enero-Diciembre</season>
<year>2020</year>
</pub-date>
<volume>33</volume>
<history>
<date date-type="received" publication-format="dd/mm/yyyy">
<day>26</day>
<month>08</month>
<year>2019</year>
</date>
<date date-type="accepted" publication-format="dd/mm/yyyy">
<day>20</day>
<month>10</month>
<year>2019</year>
</date>
<date date-type="pub" publication-format="dd/mm/yyyy">
<day>20</day>
<month>05</month>
<year>2020</year>
</date>
</history>
<permissions>
<ali:free_to_read/>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>Esta obra está bajo una Licencia Creative Commons Atribución 4.0 Internacional.</license-p>
</license>
</permissions>
<abstract xml:lang="en">
<title>Abstract</title>
<p>Energy efficiency and sustainability are important factors to address in the context of smart cities. In this sense, a necessary functionality is to reveal various preferences, behaviors, and characteristics of individual consumers, considering the energy consumption information from smart meters. In this paper, we introduce a general methodology and a specific two-level clustering approach that can be used to group, considering global and local features, energy consumptions and productions of households. Thus, characteristic load and production profiles can be determined for each consumer and prosumer, respectively. The obtained results will be generally applicable and will be useful in a general business analytics context.</p>
</abstract>
<trans-abstract xml:lang="es">
<title>Resumen</title>
<p>La eficiencia energética y la sostenibilidad son factores importantes a abordar en el contexto de las ciudades inteligentes. En este sentido, una funcionalidad necesaria consiste en revelar varias preferencias, comportamientos y características de los consumidores individuales, considerando la información de consumo de energía de los metro-contadores inteligentes. En este artículo presentamos una metodología general y un enfoque de agrupamiento en dos niveles teniendo en cuenta las características globales y locales del consumo de energía y la producción de los hogares. Por lo tanto, se pueden determinar los perfiles característicos de carga y producción para cada consumidor y prosumidor, respectivamente. Los resultados obtenidos serán de aplicación general y serán útiles en un contexto de análisis empresarial general.</p>
</trans-abstract>
<trans-abstract xml:lang="pt">
<title>Resumo</title>
<p>A eficiência energética e a sustentabilidade são fatores importantes de abordar no contexto das cidades inteligentes. Neste sentido, uma função necessária seria revelar várias preferências, comportamentos e características dos consumidores individuais, considerando a informação de consumo de energia dos medidores de dados inteligentes. Este artigo apresenta uma metodologia geral e um enfoque de agrupação em dois níveis, tendo em conta as características globais e locais do consumo de energia e a produção dos lares. Por tanto, é possível determinar os perfis característicos de carga e produção para cada consumidor e prosumidor, respetivamente. Os resultados obtidos serão de aplicação geral, especialmente, em um contexto de análise empresarial.</p>
</trans-abstract>
<kwd-group xml:lang="en">
<title>Keywords</title>
<kwd>clustering</kwd>
<kwd>time series</kwd>
<kwd>smart metering</kwd>
</kwd-group>
<kwd-group xml:lang="es">
<title>Palabras clave</title>
<kwd>agrupamiento</kwd>
<kwd>series de tiempo</kwd>
<kwd>medición inteligente</kwd>
</kwd-group>
<kwd-group xml:lang="pt">
<title>Palavras-chave</title>
<kwd>agrupação</kwd>
<kwd>séries de tempo</kwd>
<kwd>medição inteligente</kwd>
</kwd-group>
<counts>
<fig-count count="30"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="41"/>
</counts>
<custom-meta-group>
<custom-meta>
<meta-name>Cited as</meta-name>
<meta-value>Arco G., L., Casas C., G. M., &amp; Nowé A. (2020). Two-level clustering methodology for smart metering data. <italic>Cuadernos de Administración</italic>, 33. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.11144.Javeriana.cao33.tlcms">https://doi.org/10.11144.Javeriana.cao33.tlcms</ext-link>
</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<sec sec-type="intro">
<title>
<bold>Introduction</bold>
</title>
<p>On the way towards a low-carbon future, electricity networks are considered as enablers and one of the critical areas to be studied under the Strategic Energy Technologies Plan. The first European Electricity Grid Initiative –EEGI– Roadmap 2010-2018 was approved by the European Commission and the Member States alongside the creation of EEGI in June 2010. The EEGI Roadmap defines the research, development and demonstration challenges that both European transmission and distribution system operators should address in the next years with the aim to face the requirements linked to the evolution of power systems and to respond to different external factors. For this reason, smart-grid projects are receiving a lot of attention (<xref ref-type="bibr" rid="redalyc_20562876001_ref22">Hübner &amp; Prüggler, 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref20">Giordano et al., 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref28">Losa, De Nigris &amp; Van, 2013</xref>). New perspectives emerge for energy management. Many smart meters and sensors are being deployed and they result in a new data deluge we will have to face. With the rollout of smart metering infrastructure at scale, demand-response programs may now be tailored based on users’ consumption and production patterns as mined from sensed data.</p>
<p>Energy efficiency and sustainability are important factors to address in the context of smart cities. In this sense, a necessary functionality is to reveal various preferences, behaviors, and characteristics of individual consumers and prosumers, considering the fine-grained energy consumption and production information from smart meters, respectively. Smart metering and nonintrusive load monitoring play a crucial role in fighting energy thefts and for optimizing the energy consumption of the home, building, city, and so forth (<xref ref-type="bibr" rid="redalyc_20562876001_ref16">Fenza, Gallo &amp; Loia, 2019</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref3">Ahmad et al., 2018</xref>). Besides, it is very important to reduce the mismatch between the actual and expected energy demand, which is often due to an anomalous operation of the equipment and control systems. In this context, the characterization of energy consumption patterns over time is of fundamental importance (<xref ref-type="bibr" rid="redalyc_20562876001_ref10">Capozzoli et al., 2018</xref>).</p>
<p>To the best of our knowledge, all approaches are still in a research phase, especially when it comes to clustering methods consumption and production data to provide clusters for each consumer and prosumer profile (<xref ref-type="bibr" rid="redalyc_20562876001_ref21">Hossain et al., 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref7">Binh et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref31">Mutanen et al., 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref27">Lee, Haben, &amp; Grindrod, 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref6">Ardakanian et al., 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref26">Lavin &amp; Klabjan, 2014</xref>).</p>
<p>While some authors have been working on grouping consumers considering the similarity among time series models, such as ARMA and ARIMA (<xref ref-type="bibr" rid="redalyc_20562876001_ref8">Brockwell &amp; Davis, 2002</xref>); others have been focusing on grouping consumers considering the time series as feature vectors (<xref ref-type="bibr" rid="redalyc_20562876001_ref18">Flath et al., 2012</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref9">Cao, Beckel, &amp; Staake, 2013</xref>). Most of the proposals model the data following only a local point of view, others only global, and others considering the original time series, which limits the analysis. The most successful approaches have been those that combine various clustering methods (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref4">Albert &amp; Rajagopal, 2013</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref35">Räsänen et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>). It is even possible to find some proposals that combine clustering with other machine learning techniques, such as association rules (<xref ref-type="bibr" rid="redalyc_20562876001_ref19">Funde et al., 2019</xref>). Nevertheless, those hybrid approaches only exploit the combination of clustering methods in order to mitigate the disadvantages of ones and enhance the benefits of others. However, they do not exploit other important reasons for developing hybrid time series clustering models.</p>
<p>Due to the limitations expressed above, in this paper, we introduce a general methodology that combines clustering methods in two stages and exploits in a hybrid way local and global patterns of the series under analysis. Our proposal can be used to group energy consumers and prosumers according to the similarity of their daily and yearly consumption and production, respectively. Thus, characteristic load and production profiles per time period can be determined, as we will explain later. The obtained results are generally applicable and will be useful in a general business analytics context.</p>
</sec>
<sec>
<title>
<bold>Cluster analysis of smart meter data</bold>
</title>
<p>Smart meter data are time series; which makes the analysis quite complex. For that reason, cluster analysis of consumption data has been explored in some papers (<xref ref-type="bibr" rid="redalyc_20562876001_ref21">Hossain et al., 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref7">Binh et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref31">Mutanen et al., 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref27">Lee et al., 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref6">Ardakanian et al., 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref26">Lavin &amp; Klabjan, 2014</xref>), not so much the clustering of production data. From now on we will refer to consumption data clustering approaches; however, all proposals are applicable to production data clustering as well. Most authors have been focused on grouping consumers considering the time series as feature vectors. In literature four approaches are proposed to cope with the feature vector definition:</p>
<p>
<list list-type="order">
<list-item>
<p>Consider features as interval consumption measurements (e.g., every 15 minutes) (<xref ref-type="bibr" rid="redalyc_20562876001_ref18">Flath et al., 2012</xref>;<xref ref-type="bibr" rid="redalyc_20562876001_ref9"> Cao et al., 2013</xref>).</p>
</list-item>
<list-item>
<p>Only use global features (e.g., mean and standard deviation of an overall day) for characterizing each consumer (<xref ref-type="bibr" rid="redalyc_20562876001_ref26">Lavin &amp; Klabjan, 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>).</p>
</list-item>
<list-item>
<p>Extend the time series data by additional global features or other external measures (<xref ref-type="bibr" rid="redalyc_20562876001_ref6">Ardakanian et al., 2014</xref>).</p>
</list-item>
<list-item>
<p>Create local patterns for characterizing the time series (<xref ref-type="bibr" rid="redalyc_20562876001_ref27">Lee et al., 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref13">Dent et al., 2011</xref>).</p>
</list-item>
</list>
</p>
<p>The first one follows a raw-data-based approach, the last two follow a feature-based approach and the third one considers an extension of the raw data including other features.</p>
<p>The definition of a distance measure between time series is necessary for the four approaches (<xref ref-type="bibr" rid="redalyc_20562876001_ref23">Iglesias &amp; Kastner, 2013</xref>), and it depends on the clustering objective, which can be similarity in time, similarity in shape or similarity in change (<xref ref-type="bibr" rid="redalyc_20562876001_ref41">Zhang et al., 2011</xref>):</p>
<p>
<list list-type="bullet">
<list-item>
<p>The similarity in time is to cluster together series that vary in a similar way on each time level, as shown in  <xref ref-type="fig" rid="gf1">Figure 1</xref> . Usually, the clustering of time series data based on similarity in time requires the calculation of the exact distances among all the time series data in a dataset.</p>
</list-item>
<list-item>
<p>The similarity in shape is to cluster series with common shape features together, as shown in  <xref ref-type="fig" rid="gf2">Figure 2</xref> . This may constitute identifying common trends occurring at different times or similar sub-patterns in the data.</p>
</list-item>
<list-item>
<p>The similarity in change is to cluster series by the similarity in how they vary from time level to time level, as shown in <xref ref-type="fig" rid="gf3">Figure 3</xref>.</p>
</list-item>
</list>
</p>
<p>
<fig id="gf1">
<label>Figure 1</label>
<caption>
<title>Time series based on similarity in time</title>
</caption>
<alt-text>Figure 1 Time series based on similarity in time</alt-text>
<graphic xlink:href="20562876001_gf2.png" position="anchor" orientation="portrait"/>
<attrib>Source: <xref ref-type="bibr" rid="redalyc_20562876001_ref41">Zhang et al. (2011)</xref>.</attrib>
</fig>
</p>
<p>
<fig id="gf2">
<label>Figure 2</label>
<caption>
<title>Time series based on similarity in shape</title>
</caption>
<alt-text>Figure 2 Time series based on similarity in shape</alt-text>
<graphic xlink:href="20562876001_gf3.png" position="anchor" orientation="portrait"/>
<attrib>Source: <xref ref-type="bibr" rid="redalyc_20562876001_ref41">Zhang et al. (2011)</xref>.</attrib>
</fig>
</p>
<p>
<fig id="gf3">
<label>Figure 3</label>
<caption>
<title>Time series based on similarity in change</title>
</caption>
<alt-text>Figure 3 Time series based on similarity in change</alt-text>
<graphic xlink:href="20562876001_gf4.png" position="anchor" orientation="portrait"/>
<attrib>Source: <xref ref-type="bibr" rid="redalyc_20562876001_ref41">Zhang et al. (2011)</xref>.</attrib>
</fig>
</p>
<p>In this research, we are interested in time series clustering where the main clustering objective is the similarity in time because we need to cluster together series that vary in a similar way at each time interval. For this reason, in the first approach, it is necessary to define a distance measure based on the specific characteristics of time series data. Secondly, the arithmetic means of the single time segments are the starting point for the formation of global consumer behavior, but global features only do not properly represent the customers’ behavior. Thus, the second approach is not enough to segment the customers, and make groups of households with similar consumption patterns and determine on the fly the cluster membership of a given load curve. In the third approach, the dimensionality of the time series is increased and it could be difficult to manage different kinds of features, global and local in the same clustering process. Finally, the last approach could be useful for detecting clusters with similar load profiles, but it could be depending on the homogeneity of the data from the global feature point of view.</p>
<p>As we pointed out, the above approaches have some advantages and disadvantages. Thus, some authors prefer to develop hybrid approaches for time series clustering in other to solve the above disadvantages (<xref ref-type="bibr" rid="redalyc_20562876001_ref41">Zhang et al., 2011</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref25">Lai et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref2">Aghabozorgi et al., 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref40">Warren, 2007</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref33">Oates, Firoiu &amp; Cohen, 1999</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref1">Aghabozorgi, Saybani &amp; Wah, 2012</xref>). There are several reasons for developing hybrid time series clustering models (<xref ref-type="bibr" rid="redalyc_20562876001_ref25">Lai et al., 2010</xref>). For instance, we might obtain very different clustering results for the same time series dataset when different time granules are considered. For time series clustering, dimensionality reduction methods are often applied to reduce the data dimension before clustering. Consequently, the information of subsequence may be overlooked. Therefore this might result in different clustering results after considering the subsequence information. Some conventional clustering methods require prior information and domain knowledge; others do not require prior information but are too computationally expensive to be applied on very large data sets. The combination of clustering methods can mitigate the disadvantages of some and enhance the benefits of others. For some applications, the clustering objective might not be that apparent. The selection of the time series representation and the similarity measure depends on the clustering objective. Thus, different clustering approaches are required.</p>
<p>Some hybrid clustering methods are proposed in the area of clustering analysis of smart metering data (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>;<xref ref-type="bibr" rid="redalyc_20562876001_ref4"> Albert &amp; Rajagopal, 2013</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref35">Räsänen et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>). Most of them apply Self-Organizing Maps –SOM– (<xref ref-type="bibr" rid="redalyc_20562876001_ref24">Kohonen, 1982</xref>) in the first level and k-means (<xref ref-type="bibr" rid="redalyc_20562876001_ref29">MacQueen, 1967</xref>) or hierarchical clustering algorithms (<xref ref-type="bibr" rid="redalyc_20562876001_ref35">Räsänen et al., 2010</xref>) in the second level. SOM is used to obtain a reduction of the dimension of the initial dataset and k-means is used to group the weight vectors of the SOM’s units and the final clusters are obtained (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2015</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref35">Räsänen et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>). Another approach applies k-means first and uses spectral clustering to segment a collection into classes of similar statistical properties (<xref ref-type="bibr" rid="redalyc_20562876001_ref4">Albert &amp; Rajagopal, 2013</xref>). These hybrid approaches only exploit the combination of clustering methods in order to mitigate the disadvantages of ones and enhance the benefits of others. However, they do not exploit other important reasons for developing hybrid time series clustering models.</p>
</sec>
<sec>
<title>
<bold>General ideas, stages, and schema of the proposed methodology</bold>
</title>
<p>We introduce a general methodology that can be used to group time series considering different time granules. The proposed methodology consists of the following stages, as shown in <xref ref-type="fig" rid="gf4">Figure 4</xref>.</p>
<p>
<fig id="gf4">
<label>Figure 4</label>
<caption>
<title>General schema of the proposed methodology</title>
</caption>
<alt-text>Figure 4 General schema of the proposed methodology</alt-text>
<graphic xlink:href="20562876001_gf5.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<bold> Stage 1: Data gathering.</bold> Read time series; e.g., the time intervals of interest for data representation are typical 1 min, 15 min or 1h in the context of smart metering applications (<xref ref-type="bibr" rid="redalyc_20562876001_ref11">Chicco, 2012</xref>).</p>
<p>
<bold> Stage 2: Data stratifying.</bold> Segment the data, which separates the raw data sets into more homogeneous subsets, in order to sustain scalability; e.g., data can be stratified using a split between weekend and weekdays, or between summer and winter months when we are working with smart metering data (<xref ref-type="bibr" rid="redalyc_20562876001_ref18">Flath et al., 2012</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref9">Cao et al., 2013</xref>).</p>
<p>
<bold> Stage 3: Data cleaning.</bold> Detect and remove errors and inconsistencies from data in order to improve the data quality. In the smart metering domain some strategies can discard data sets showing more than one hour of recording gaps (<xref ref-type="bibr" rid="redalyc_20562876001_ref18">Flath et al., 2012</xref>); check for inconsistencies in the data and remove outliers (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>); detect missing values and replace them using regression techniques (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>); remove special days (e.g., public holidays) (<xref ref-type="bibr" rid="redalyc_20562876001_ref9">Cao et al., 2013</xref>); remove non-continuous data (<xref ref-type="bibr" rid="redalyc_20562876001_ref30">McLoughlin, Duffy &amp; Conlon, 2012</xref>).</p>
<p>
<bold> Stage 4: Different time granule clustering.</bold> Select the time granule before clustering. Depending on the granularity level desired in the clustering, it is defined all the elements that contribute to the clustering. This stage can be repeated several times depending on how much you want to refine the level of granularity in the data analysis. This is the most important stage of our methodology. For that reason, we will explain in detail its main steps:</p>
<p>
<list list-type="simple">
<list-item>
<p>4.1 Time granule selection</p>
</list-item>
<list-item>
<p>4.2 Data representation</p>
</list-item>
<list-item>
<p>4.3 Data preprocessing</p>
</list-item>
<list-item>
<p>4.4 Distance/similarity selection</p>
</list-item>
<list-item>
<p>4.5 Clustering algorithm selection</p>
</list-item>
</list>
</p>
<p>
<bold> Stage 5: Post-clustering.</bold> Apply clustering validation techniques, visualize the clustering results and obtain labels and prototypes for each cluster. For example, a useful post-clustering result in the smart metering applications can be the calculation of the global power and energy information for the customer classes for tariff setting purposes (<xref ref-type="bibr" rid="redalyc_20562876001_ref11">Chicco, 2012</xref>).</p>
<p>The selection of the level of granularity is closely associated with the objective of the desired clustering, as shown in <xref ref-type="fig" rid="gf5">Figure 5</xref>. In this step, it is necessary to decide if the objective is the similarity in shape, in change or in time. The selection of the time series representation (raw-data-based, feature-based or model-based representation) also depends on the clustering objective. For example, if the objective is the similarity in time then we suggest using raw-data-based representation. On the other hand, if the objective is the similarity in change, we suggest using a model-based representation. We can change the representation in different clustering levels.</p>
<p>
<fig id="gf5">
<label>Figure 5</label>
<caption>
<title>Stage 4 schema</title>
</caption>
<alt-text>Figure 5 Stage 4 schema</alt-text>
<graphic xlink:href="20562876001_gf6.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>Then, preprocessing is in charge of applying dimensionality reduction and normalization methods in correspondence with the clustering objective and the selected representation. Finally, it is necessary to define the appropriate distance measure and apply a clustering algorithm for obtaining clusters where all the series grouped in the same cluster should be coherent or homogeneous. It is important to take into account which is our clustering objective for deciding which distance measure we will apply. For example, if the clustering objective is the similarity in shape, then it will be useful to apply Dynamic Time Warping –DTW– distance. The most used algorithms are k-means (<xref ref-type="bibr" rid="redalyc_20562876001_ref26">Lavin &amp; Klabjan, 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref18">Flath et al., 2012</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>), Self-Organizing Maps –SOM– (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref30">McLoughlin et al., 2012</xref>), hierarchical approaches (<xref ref-type="bibr" rid="redalyc_20562876001_ref9">Cao et al., 2013</xref>) and Expectation Maximization (<xref ref-type="bibr" rid="redalyc_20562876001_ref4">Albert &amp; Rajagopal, 2013</xref>); as well as hybrid approaches (<xref ref-type="bibr" rid="redalyc_20562876001_ref17">Figueiredo et al., 2005</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref4">Albert &amp; Rajagopal, 2013</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref35">Räsänen et al., 2010</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>).</p>
<p>Stage 4 offers the possibility to design procedures for different clustering objectives using diverse granules in the time series representation. Taking into account the smart metering domain, it could be useful in the first clustering level to group consumption or production data considering their voltage level combined with general global features; thus, a normalization process, in the second level, could be carried out inclusive a similarity in time objective clustering. In the case of high dimensional series, it could be useful to apply a feature-based representation in the first clustering level, and after that, refine the clustering results considering the raw-representation in the second level.</p>
</sec>
<sec>
<title>
<bold>Two-level clustering approach based on local and global features</bold>
</title>
<p>In this section, we apply the general architecture outlined in the previous section, to the smart grid data introduced earlier. More specifically, we apply a two-level clustering approach. The first level of clustering is based on features extracted from the time series. Regardless of the length of the time series and missing values, a finite set of statistical measures is used to capture the nature of the time series. The feature values are obtained from each individual series and can be fed into some specific clustering algorithm. In the first level, we propose to cluster data using only global features in order to divide consumers or prosumers considering their daily consumption or production data, respectively. In the second level, we split the obtained clusters in the first level, considering local features for discovering sub-clusters for each consumption or production profile. Features are obtained by applying statistical operations that best capture the underlying characteristics of the time series, depending on the clustering objective. Thus, a feature-based representation is used at both levels; nevertheless, the clustering objective is different in each level. The main objective is clustering considering the similarity in time.</p>
<p>Transforming the raw time-series data into the set of features has been used by several authors (<xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref38">Wang, Smith &amp; Hyndman, 2006</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref39">Wang et al., 2004</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref32">Nanopoulos, Alcock, &amp; Manolopoulos, 2001</xref>). Feature-based representation has several advantages; we will mention some of them. When the time series is very long (high dimensionality), some clustering algorithms become intractable; for instance, fail because the similarity is dubious in high dimension space. Applying dimensionality reduction via feature extraction, we are able to cluster long length time series very efficiently. Despite the length of the time series and missing values, a finite set of statistical measures can be used to capture the global and local nature of the time series. Furthermore, feature extraction is used to compress large data sets by means of dimensionality reduction. When the clustering algorithm is based on a distance metric (e.g., Euclidean distance), it cannot handle time series with missing data or of different lengths if actual points are used as inputs. However, by extracting a set of measures from the original time series we simply bypass this problem. Computational efficiency can be increased and the use of more sophisticated clustering algorithms is possible. When the clustering objective is the similarity in shape, it is possible to obtain good results using a feature-based representation.</p>
<p>Nevertheless, feature-based representation has some disadvantages; we will mention some of them. The majority of feature extraction methods are generic in nature, the extracted features are usually application dependent. Thus one set of features that work well on one application might not be relevant to another. When the clustering objective is the similarity in time, it is not possible to obtain good results using global features extracted from the time series.</p>
<p>Considering the above-mentioned advantages and disadvantages, and bearing in mind the objective of the clustering at each level, we propose to obtain two kinds of features at each level. <xref ref-type="fig" rid="gf6">Figure 6</xref> shows the general schema of the proposed two-level clustering approach based on local and global features.</p>
<p>
<fig id="gf6">
<label>
<bold>Figure 6</bold>
</label>
<caption>
<title>General schema of our two-level clustering approach based on local and global features</title>
</caption>
<alt-text>Figure 6 General schema of our two-level clustering approach based on local and global features</alt-text>
<graphic xlink:href="20562876001_gf7.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>In the first level, we propose to cluster data using only global features in order to divide consumptions or productions considering general behaviors. The principal global features to compute are mean, minimum, maximum and sum considering the total original features (e.g., 96 features if we consider 15 min interval consumption or production during a day). The clustering results can be improved if we include other features such as median, mode, standard deviation, variance, skewness, kurtosis, range, trend, seasonality, periodicity, serial correlation and chaos (<xref ref-type="bibr" rid="redalyc_20562876001_ref38">Wang et al., 2006</xref>). Using a global feature-based representation, the dimensionality of the time series is significantly reduced and the clustering algorithm is much less sensitive to missing or noisy data. These features are enough to obtain clusters of consumers or prosumers with the same consumption or production levels and general characteristics, respectively; but they are not enough for generating clusters for each consumption or production profile per time period, because they cannot identify peaks at specific time periods.</p>
<p>In the second level, we split and refine the obtained clusters at the first level, considering local features for discovering sub-clusters. The local features express other information than the global features and emphasize the original time series characteristics. For characterizing the time series locally, it is possible to define a one hour, one week, or one day window, depending on the original time period. If we are processing daily profiles, a one hour window is used. A one day window is used for processing yearly profiles. We created four features for each window, computing mean, maximum, minimum and range, respectively. These local values allow expressing the time series behavior in each specified window. The clustering results can be improved if we include other features such as median, mode, standard deviation, variance, skewness, kurtosis, range, trend, seasonality, periodicity, serial correlation and chaos (<xref ref-type="bibr" rid="redalyc_20562876001_ref34">Räsänen &amp; Kolehmainen, 2009</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref38">Wang et al., 2006</xref>).</p>
<p>We currently explore suitable clustering methods and the choice of parameters that enable us to obtain general and specific clusters considering global and local features, respectively. The most used clustering methods are k-means (<xref ref-type="bibr" rid="redalyc_20562876001_ref29">MacQueen, 1967</xref>), linkage (<xref ref-type="bibr" rid="redalyc_20562876001_ref12">Defays, 1977</xref>), spectral clustering (<xref ref-type="bibr" rid="redalyc_20562876001_ref37">Shi &amp; Malik, 2000</xref>) and SOM (<xref ref-type="bibr" rid="redalyc_20562876001_ref24">Kohonen, 1982</xref>). Useful similarity measures for these algorithms are Euclidean, cosine, correlation, and Manhattan (<xref ref-type="bibr" rid="redalyc_20562876001_ref23">Iglesias &amp; Kastner, 2013</xref>). We applied the k-means clustering algorithm (<xref ref-type="bibr" rid="redalyc_20562876001_ref29">MacQueen, 1967</xref>) on global features in the first level for fixing a reasonable number of clusters to be obtained. The Complete Linkage clustering algorithm (<xref ref-type="bibr" rid="redalyc_20562876001_ref12">Defays, 1977</xref>) was applied to the second level. The Euclidean distance was used for comparing the vectors in the cluster because it is a good distance when the clustering objective is the similarity in time.</p>
</sec>
<sec>
<title>
<bold>A study case on belgian data</bold>
</title>
<p>In Belgium, the authority on energy policy is shared between the federal and the regional administrations. The competent authority for the smart metering roll-out in Flanders is the regional energy regulator, VREG, while there are two Distribution System Operators –DSO–: Eandis and Infrax (<xref ref-type="bibr" rid="redalyc_20562876001_ref15">European Commission, 2014</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref36">Renner &amp; Heinemann, 2011</xref>).</p>
<p>There are few studies on profile identification and consumer segmentation in Belgium (<xref ref-type="bibr" rid="redalyc_20562876001_ref14">Espinoza et al., 2005</xref>; <xref ref-type="bibr" rid="redalyc_20562876001_ref5">Alzate et al., 2009</xref>). The least recent result starts from consumption data containing hourly consumption values from substations within the Belgian grid. The typical daily profile for each consumer is first identified, and, after that, the k-means algorithm is applied for capturing the different profiles. A large dataset of over 1300 load profiles of residential customers forms the basis for modeling in the most recent result. Each load profile is a sequence of measured data, with a resolution of 15 min, over the duration of one year. A multiway spectral clustering without the use of pre-modeling steps was used to detect consumer profiles.</p>
<p>In the research presented in this paper, we use real-life data, provided by Eandis. Houses in Belgium are connected to the electricity grid of the DSO and they are organized in neighborhoods of different sizes. All houses in a given neighborhood are connected via the low-voltage grid to one substation of the DSO. The dataset is comprised of aggregated 15 min intervals of electricity consumption and production of 2928 homes from 44 substations in Belgium.</p>
<p>Our objective is to apply the proposed methodology, specifically the two-level clustering algorithm, to detect the consumption and production profiles considering the customer behaviors in order to contribute to future decision-making problems in this field.</p>
<sec>
<title>
<bold>
<italic>Consumer and prosumer data</italic>
</bold>
</title>
<p>The structure of a consumer or prosumer data is provided in a table, where each row represents a consumption or production day of one consumer or prosumer, respectively, and each numbered column represents a 15 min consumption or production interval, respectively.</p>
<p>The consumption behaviors change depending on weekends or weekdays. <xref ref-type="fig" rid="gf7">Figure 7</xref> shows the daily consumption of a particular house for one week. Since there is a lot of variability over the different days it is not possible to identify an overall consumer profile. <xref ref-type="fig" rid="gf8">Figure 8</xref> illustrates different behaviors in a specific weekend for one consumer. Notice that it is not possible to detect a weekend profile for this consumer. <xref ref-type="fig" rid="gf9">Figure 9</xref> shows the daily consumption series from Monday to Friday for the same consumer. The profiles for weekdays are clearer than for the weekend considering this particular example. Thus, we are interested in the detection of consumption and production profiles, these profiles do not necessarily coincide with the consumers and prosumers profiles.</p>
<p>
<fig id="gf7">
<label>Figure 7</label>
<caption>
<title>Daily consumption of a house in a whole week</title>
</caption>
<alt-text>Figure 7 Daily consumption of a house in a whole week</alt-text>
<graphic xlink:href="20562876001_gf8.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf8">
<label>Figure 8</label>
<caption>
<title>Daily consumption of a house in a weekend</title>
</caption>
<alt-text>Figure 8 Daily consumption of a house in a weekend</alt-text>
<graphic xlink:href="20562876001_gf9.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf9">
<label>Figure 9</label>
<caption>
<title>Daily consumption of a house for weekends</title>
</caption>
<alt-text>Figure 9 Daily consumption of a house for weekends</alt-text>
<graphic xlink:href="20562876001_gf10.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<xref ref-type="fig" rid="gf7">Figure 7</xref>, <xref ref-type="fig" rid="gf8">Figure 8</xref> and <xref ref-type="fig" rid="gf9">Figure 9</xref> suggest the necessity of segmentation of the analysis considering weekends and weekdays separately. This is one criterion for splitting the raw data sets into more homogeneous subsets, besides it supports scalability. The seasonality effect also influences the shape of the profiles.</p>
</sec>
<sec>
<title>
<bold>
<italic>Experimental results</italic>
</bold>
</title>
<p>We applied the two-level clustering approach based on local and global features to the dataset with the electricity consumption and production of 2928 homes from 44 substations in Belgium, considering the period between 1<sup>st</sup> of November 2013 and 31<sup>st</sup> of October 2014.</p>
<p>Discovering typical daily consumption and production profiles is the main objective. Further, we want to discover the most typical consumers and prosumers considering their yearly consumption and production behavior, respectively. To this extent, four experiments were designed.</p>
</sec>
<sec>
<title>
<italic>
<bold>Daily consumption and production profiles</bold>
</italic>
</title>
<p>We prepared two datasets for discovering daily profiles. One with daily consumption and the other with the daily production values. Each time series consists of 96 energy consumption or production intervals. The global features were obtained by computing the mean, minimum, maximum, sum, median, standard deviation, variance and range considering the 96 original features for each daily consumption or production vector. For the second level, we created local features in order to refine the initial clustering results. The local features have to express more information than the original time series and global features. Thus, we created four features for each hour (i.e., four 15 min intervals), computing the mean, maximum, minimum and range.</p>
<p>The two-level clustering algorithm was applied to both daily consumption and production data. We will present here the obtained results working with the daily consumption values.</p>
<p>
<fig id="gf10">
<label>Figure 10</label>
<caption>
<title>Cluster 16 from the first level</title>
</caption>
<alt-text>Figure 10 Cluster 16 from the first level</alt-text>
<graphic xlink:href="20562876001_gf11.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf11">
<label>Figure 11</label>
<caption>
<title>Cluster 28 from the first level</title>
</caption>
<alt-text>Figure 11 Cluster 28 from the first level</alt-text>
<graphic xlink:href="20562876001_gf12.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>Figures <xref ref-type="fig" rid="gf10">10</xref> and <xref ref-type="fig" rid="gf11">11</xref> show a selection of clusters obtained in the first level. These are the only two clusters from 38 obtained clusters in the first level, which express typical profiles only considering global features.</p>
<p>The clusters showed in <xref ref-type="fig" rid="gf12">Figure 12</xref>, <xref ref-type="fig" rid="gf13">Figure 13</xref> and <xref ref-type="fig" rid="gf14">Figure 14</xref> represent the clusters containing the majority of the time series. They illustrate the necessity of applying the second level for refining the clustering results and consequently obtain the daily consumption profiles.</p>
<p>
<fig id="gf12">
<label>Figure 12</label>
<caption>
<title>Cluster 10 from the first level</title>
</caption>
<alt-text>Figure 12 Cluster 10 from the first level</alt-text>
<graphic xlink:href="20562876001_gf13.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf13">
<label>Figure 13</label>
<caption>
<title>Cluster 15 from the first level</title>
</caption>
<alt-text>Figure 13 Cluster 15 from the first level</alt-text>
<graphic xlink:href="20562876001_gf15.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf14">
<label>Figure 14</label>
<caption>
<title>Cluster 31 from the first level</title>
</caption>
<alt-text>Figure 14 Cluster 31 from the first level</alt-text>
<graphic xlink:href="20562876001_gf16.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>After applying the second level, we really obtained the most representative daily consumption prototypes. We classified the most representative prototypes in the following categories: high, medium and low consumption prototypes. This classification is possible by only considering the first level results. Nevertheless, we refine this classification on the second level considering the consumption peaks and daily consumption behavior as we show in the following <xref ref-type="fig" rid="gf15">figures</xref>.</p>
<p>
<fig id="gf15">
<label>Figure 15</label>
<caption>
<title>High daily consumption prototype</title>
</caption>
<alt-text>Figure 15 High daily consumption prototype</alt-text>
<graphic xlink:href="20562876001_gf17.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf16">
<label>Figure 16</label>
<caption>
<title>Medium consumption prototype with higher consumption in the morning</title>
</caption>
<alt-text>Figure 16 Medium consumption prototype with higher consumption in the morning</alt-text>
<graphic xlink:href="20562876001_gf18.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf17">
<label>Figure 17</label>
<caption>
<title>Medium consumption prototype with higher consumption in the afternoon</title>
</caption>
<alt-text>Figure 17 Medium consumption prototype with higher consumption in the afternoon</alt-text>
<graphic xlink:href="20562876001_gf19.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>We obtained six clusters that represent high consumption prototypes. Figure 15 shows one of them.</p>
<p>
<fig id="gf18">
<label>Figure 18</label>
<caption>
<title>Medium consumption prototypes with higher consumption at afternoon-night</title>
</caption>
<alt-text>Figure 18 Medium consumption prototypes with higher consumption at afternoon-night</alt-text>
<graphic xlink:href="20562876001_gf20.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>Two clusters representing medium consumption prototypes with the higher consumption in the morning were obtained in the second level. <xref ref-type="fig" rid="gf16">Figure 16</xref> shows one of them. <xref ref-type="fig" rid="gf17">Figure 17</xref> shows one of the four obtained medium consumption prototypes with higher consumption in the afternoon; whereas <xref ref-type="fig" rid="gf18">Figure 18</xref> shows one of the 10 detected medium consumption prototypes with consumption peaks at afternoon-night.</p>
<p>We identified four clusters that represent medium consumption prototypes with small consumption peaks in the morning and high consumption peaks in the afternoon, <xref ref-type="fig" rid="gf19">Figure 19</xref> shows one of them.</p>
<p>
<fig id="gf19">
<label>Figure 19</label>
<caption>
<title>Medium consumption prototype with small consumption peaks in the morning and high consumption peaks in the afternoon</title>
</caption>
<alt-text>Figure 19 Medium consumption prototype with small consumption peaks in the morning and high consumption peaks in the afternoon</alt-text>
<graphic xlink:href="20562876001_gf21.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>Some consumers have a medium consumption from morning to afternoon, which is more or less constant. <xref ref-type="fig" rid="gf20">Figure 20</xref> shows this kind of consumer.</p>
<p>
<fig id="gf20">
<label>Figure 20</label>
<caption>
<title>Typical consumption prototype with stable medium consumption from morning to afternoon</title>
</caption>
<alt-text>Figure 20 Typical consumption prototype with stable medium consumption from morning to afternoon</alt-text>
<graphic xlink:href="20562876001_gf22.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>The majority of the consumers have consumption peaks in the morning, in the afternoon and at night. Twelve typical medium consumption prototypes with these three peaks were found. <xref ref-type="fig" rid="gf21">Figure 21</xref> shows one of them.</p>
<p>
<fig id="gf21">
<label>Figure 21</label>
<caption>
<title>Typical medium consumption prototype with three consumption peaks in a day</title>
</caption>
<alt-text>Figure 21 Typical medium consumption prototype with three consumption peaks in a day</alt-text>
<graphic xlink:href="20562876001_gf23.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>Other consumers have a more variable consumption pattern because they have a lot of consumption peaks during the day. <xref ref-type="fig" rid="gf22">Figure 22</xref> shows one of these kinds of daily consumption prototypes.</p>
<p>
<fig id="gf22">
<label>Figure 22</label>
<caption>
<title>Typical medium consumption prototype with many consumption peaks in a day</title>
</caption>
<alt-text>Figure 22 Typical medium consumption prototype with many consumption peaks in a day</alt-text>
<graphic xlink:href="20562876001_gf24.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>We obtained two categories for grouping the low consumption prototypes, one of them with a stable consumption and the other one with peaks of consumption. <xref ref-type="fig" rid="gf23">Figure 23</xref> and <xref ref-type="fig" rid="gf25">Figure 24</xref> show representative clusters of these categories, respectively.</p>
<p>
<fig id="gf23">
<label>Figure 23</label>
<caption>
<title>Low and stable consumption prototype</title>
</caption>
<alt-text>Figure 23 Low and stable consumption prototype</alt-text>
<graphic xlink:href="20562876001_gf25.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf25">
<label>Figure 24</label>
<caption>
<title>Low and unstable consumption prototype</title>
</caption>
<alt-text>Figure 24 Low and unstable consumption prototype</alt-text>
<graphic xlink:href="20562876001_gf27.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>The two-level clustering approach allows us to discover prototypes considering global consumption levels and consumption behaviors. We obtain clusters with homogeneous consumption levels in the first level and clusters with the same profile in the second level. However, two clusters obtained in the second level can look similar considering the consumption behavior, but they are very different taking into account the consumption levels. For instance, prototypes shown in Figure 25(<xref ref-type="fig" rid="gf26">a</xref>, <xref ref-type="fig" rid="gf27">b</xref>) have similar behavior but their general consumptions are different, between 0.05 kWh and 1.04 kWh in the top one, and between 0.00 kWh and 0.23 kWh for the bottom one.</p>
<p>
<fig id="gf26">
<label>Figure 25a</label>
<caption>
<title>Two clusters with similar profiles and different consumption levels</title>
</caption>
<alt-text>Figure 25a Two clusters with similar profiles and different consumption levels</alt-text>
<graphic xlink:href="20562876001_gf28.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf27">
<label>Figure 25b</label>
<caption>
<title>Two clusters with similar profiles and different consumption levels</title>
</caption>
<alt-text>Figure 25b Two clusters with similar profiles and different consumption levels</alt-text>
<graphic xlink:href="20562876001_gf29.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
</sec>
<sec>
<title>
<bold>
<italic>Yearly consumption and production profiles</italic>
</bold>
</title>
<p>As we proceeded with the daily consumption and production data, we prepared two datasets for discovering yearly profiles. One with yearly consumption values and the other one with yearly production values. In both cases, each time series consists of 35040 energy consumption or production intervals, between 1st of November 2013 and 31<sup>st</sup> of October 2014. The global features were obtained by computing the mean, minimum, maximum, sum, median, variance, standard deviation and range considering the 35040 original features for each yearly consumption or production vector. For the second level, we created local features in order to refine the initial clustering results. In this case, the local features express the daily consumption or production, because we created eight features summarizing the information of each set of 96 consumption or production intervals by computing the mean, minimum, maximum, sum, median, standard deviation, variance and range statistics. Thus, we will not consider explicitly the daily consumption or productions behaviors.</p>
<p>The two-level clustering algorithm was applied to both yearly consumption and production matrices. Below we present the obtained results working with the yearly production values of 375 prosumers.</p>
<p>
<fig id="gf28">
<label>Figure 26</label>
<caption>
<title>Yearly production profile and its most representative prosumer (35022538)</title>
</caption>
<alt-text>Figure 26 Yearly production profile and its most representative prosumer (35022538)</alt-text>
<graphic xlink:href="20562876001_gf30.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf29">
<label>Figure 27</label>
<caption>
<title>Yearly production profile and its most representative prosumer (35022611)</title>
</caption>
<alt-text>Figure 27 Yearly production profile and its most representative prosumer (35022611)</alt-text>
<graphic xlink:href="20562876001_gf31.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>
<fig id="gf30">
<label>Figure 28</label>
<caption>
<title>Yearly production profile and its most representative prosumer (3187965)</title>
</caption>
<alt-text>Figure 28 Yearly production profile and its most representative prosumer (3187965)</alt-text>
<graphic xlink:href="20562876001_gf32.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
<p>As we can see in the following figures, the yearly production behavior is similar; nevertheless, we can see clusters with different production levels. Prosumers increase production between March and September. <xref ref-type="fig" rid="gf28">Figure 26 </xref>shows a cluster with the lowest yearly productions and <xref ref-type="fig" rid="gf29">Figure 27</xref> shows a cluster with the highest yearly productions.</p>
<p>
<fig id="gf31">
<label>Figure 29</label>
<caption>
<title>Yearly production profile and its most representative prosumer (3188111)</title>
</caption>
<alt-text>Figure 29 Yearly production profile and its most representative prosumer (3188111)</alt-text>
<graphic xlink:href="20562876001_gf33.png" position="anchor" orientation="portrait"/>
<attrib>Source: Own elaboration.</attrib>
</fig>
</p>
</sec>
</sec>
<sec sec-type="conclusions">
<title>
<bold>Conclusions and future work</bold>
</title>
<p>The proposed methodology was used through the application of the two-level clustering approach to Belgian energy consumption and production data. Daily consumption and production profiles were obtained, considering global features such as total daily consumption or production, and local features such as hourly consumption or production, respectively. Whereas, prototypical consumers and prosumers were discovered considering global features such as total yearly consumption or production, and local features such as daily consumption or production, respectively.</p>
<p>The proposed methodology has the following advantages. It allows obtaining different clustering results by using different time granules. Moreover, it allows considering different clustering objectives. In the first level, only some general statistics about the data are required; while in the second level, the main objective is the similarity in time for identifying the consumption or production profiles. The approach also allows for dimensionality reduction via feature extraction in the first level; doing so the clustering algorithm becomes more efficient. It extracts a set of measures from the original time series; as such it is possible to obtain good results using the Euclidean distance, whereas this measure cannot handle the original time series directly. It uses a finite set of statistical measures to capture the global and local nature of the time series; thus, the computational efficiency of the clustering algorithms can be improved and the use of more advanced clustering algorithms becomes possible. Finally, our approach allows obtaining clusters with homogeneous consumption or production levels in the first level and clusters with the same profile in the second level. We believe this latter characteristic is appealing to test decision policies as it is crucial that the situations are considered to consist of typical situations that might occur in reality.</p>
</sec>
</body>
<back>
<ack>
	<title>Acknowledgement</title>
	<p>The research we present in this paper was performed in the framework of the Erasmus Mundus Action 2 and in the context of the SCANERGY project. Erasmus Mundus Action 2, under the terms of the grant agreement number 2013-2591, EUREKA SD, finances a full time stay at the Postdoctoral level in the faculty of Sciences at Vrije Universiteit Brussel. SCANERGY project has received funding from the European Union's 7<sup>th</sup> Frame Program for research, technological development and demonstration under grant agreement number 324321.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="redalyc_20562876001_ref1">
<mixed-citation>Aghabozorgi, S., Saybani, M., &amp; Wah, T. (2012). Incremental clustering of time-series by fuzzy clustering. <italic>Journal of Information Science and Engineering, 28</italic>, 671-688. <ext-link ext-link-type="uri" xlink:href="https://www.iis.sinica.edu.tw/page/jise/2012/201207_03.pdf">https://www.iis.sinica.edu.tw/page/jise/2012/201207_03.pdf</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aghabozorgi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Saybani</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wah</surname>
<given-names>T.</given-names>
</name>
</person-group>
<article-title>Incremental clustering of time-series by fuzzy clustering</article-title>
<source>Journal of Information Science and Engineering</source>
<year>2012</year>
<volume>28</volume>
<fpage>671</fpage>
<lpage>688</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://www.iis.sinica.edu.tw/page/jise/2012/201207_03.pdf">https://www.iis.sinica.edu.tw/page/jise/2012/201207_03.pdf</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref2">
<mixed-citation>Aghabozorgi, S., Ying Wah, T., Herawan, T., Jalab, H., Shaygan, M., &amp; Jalali, A. (2014). A hybrid algorithm for clustering of time series data based on affinity search technique. <italic>Scientific World Journal, 2014</italic>, 1301-1314. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1155/2014/562194">https://doi.org/10.1155/2014/562194</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aghabozorgi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ying Wah</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Herawan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Jalab</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shaygan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jalali</surname>
<given-names>A.</given-names>
</name>
</person-group>
<article-title>A hybrid algorithm for clustering of time series data based on affinity search technique</article-title>
<source>Scientific World Journal</source>
<year>2014</year>
<volume>2014</volume>
<fpage>1301</fpage>
<lpage>1314</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1155/2014/562194">https://doi.org/10.1155/2014/562194</ext-link>
</comment>
<pub-id pub-id-type="art-access-id">10.1155/2014/562194</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref3">
<mixed-citation>Ahmad, T., Chen, H., Wang, J., &amp; Guo, Y. (2018). Review of various modeling techniques for the detection of electricity theft in smart grid environment. <italic>Renewable &amp; Sustainable Energy Review, 82</italic>(August), 2916-2933. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.rser.2017.10.040">https://doi.org/10.1016/j.rser.2017.10.040</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ahmad</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
</person-group>
<article-title>Review of various modeling techniques for the detection of electricity theft in smart grid environment</article-title>
<source>Renewable &amp; Sustainable Energy Review</source>
<year>2018</year>
<volume>82</volume>
<issue>August</issue>
<fpage>2916</fpage>
<lpage>2933</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.rser.2017.10.040">https://doi.org/10.1016/j.rser.2017.10.040</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.rser.2017.10.040</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref4">
<mixed-citation>Albert, A., &amp; Rajagopal, R. (2013). Smart meter driven segmentation: What your consumption says about you. <italic>IEEE Transaction. Power Systems, 28</italic>, 4019-4030. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRS.2013.2266122">https://doi.org/10.1109/TPWRS.2013.2266122</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Albert</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rajagopal</surname>
<given-names>R.</given-names>
</name>
</person-group>
<article-title>Smart meter driven segmentation: What your consumption says about you</article-title>
<source>IEEE Transaction Power Systems</source>
<year>2013</year>
<volume>28</volume>
<fpage>4019</fpage>
<lpage>4030</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRS.2013.2266122">https://doi.org/10.1109/TPWRS.2013.2266122</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/TPWRS.2013.2266122</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref5">
<mixed-citation>Alzate, C., Espinoza, M., De Moor, B., &amp; Suykens, J. (2009). Identifying customer profiles in power load time series using spectral clustering. <italic>Artificial Neural Networks – ICANN</italic>, 2009, <italic>LNCS 5769</italic>, 315-324. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-642-04277-5_32">https://doi.org/10.1007/978-3-642-04277-5_32</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alzate</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Espinoza</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>De Moor</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Suykens</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>Identifying customer profiles in power load time series using spectral clustering</article-title>
<source>Artificial Neural Networks – ICANN</source>
<year>2009</year>
<fpage>315</fpage>
<lpage>324</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-642-04277-5_32">https://doi.org/10.1007/978-3-642-04277-5_32</ext-link>
</comment>
<comment>LNCS 5769</comment>
<pub-id pub-id-type="doi">10.1007/978-3-642-04277-5_32</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref6">
<mixed-citation>Ardakanian, O., Koochakzadeh, N., Singh, R., Golab, L., &amp; Keshav, S. (2014). Computing electricity consumption profiles from household smart meter data. In EDBT Workshop on Energy Data Management (pp. 140-147). <ext-link ext-link-type="uri" xlink:href="https://pdfs.semanticscholar.org/11b9/5c1d7861e7932919394f65487b551ab3e1cd.pdf">https://pdfs.semanticscholar.org/11b9/5c1d7861e7932919394f65487b551ab3e1cd.pdf</ext-link>
</mixed-citation>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Ardakanian</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Koochakzadeh</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Golab</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Keshav</surname>
<given-names>S.</given-names>
</name>
</person-group>
<article-title>Computing electricity consumption profiles from household smart meter data</article-title>
<source>EDBT Workshop on Energy Data Management</source>
<year>2014</year>
<fpage>140</fpage>
<lpage>147</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://pdfs.semanticscholar.org/11b9/5c1d7861e7932919394f65487b551ab3e1cd.pdf">https://pdfs.semanticscholar.org/11b9/5c1d7861e7932919394f65487b551ab3e1cd.pdf</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref7">
<mixed-citation>Binh, P., Ha, N., Tuan, T., &amp; Khoa, L. (2010). Determination of representative load curve based on Fuzzy K-Means. 4th International Power Engineering and Optimization Conference (PEOCO), Shah Alam, pp. 281-286. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/PEOCO.2010.5559257">https://doi.org/10.1109/PEOCO.2010.5559257</ext-link>
</mixed-citation>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Binh</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ha</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Tuan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Khoa</surname>
<given-names>L.</given-names>
</name>
</person-group>
<source>4th International Power Engineering and Optimization Conference (PEOCO)</source>
<year>2010</year>
<fpage>281</fpage>
<lpage>286</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/PEOCO.2010.5559257">https://doi.org/10.1109/PEOCO.2010.5559257</ext-link>
</comment>
<pub-id pub-id-type="doi">0.1109/PEOCO.2010.5559257</pub-id>
<conf-name>Determination of representative load curve based on Fuzzy K-Means</conf-name>
<conf-loc>Shah Alam</conf-loc>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref8">
<mixed-citation>Brockwell, P., &amp; Davis, R. (2002). <italic>Introduction to time series and forecasting</italic>, 2 ed. Springer Texts in Statistics. New York: Springer Verlag.</mixed-citation>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Brockwell</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Davis</surname>
<given-names>R.</given-names>
</name>
</person-group>
<source>Introduction to time series and forecasting</source>
<year>2002</year>
<publisher-loc>Springer Texts in Statistics. New York</publisher-loc>
<publisher-name>Springer Verlag</publisher-name>
<edition>2 ed.</edition>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref9">
<mixed-citation>Cao, H., Beckel, C., &amp; Staake, T. (2013). Are domestic load profiles stable over time? An attempt to identify target households for demand side management campaigns, in 39th Annual Conference of the IEEE Industrial Electronics Society (pp. 4733-4738). IECON. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/IECON.2013.6699900">https://doi.org/10.1109/IECON.2013.6699900</ext-link>
</mixed-citation>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Beckel</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Staake</surname>
<given-names>T.</given-names>
</name>
</person-group>
<source>39th Annual Conference of the IEEE Industrial Electronics Society</source>
<year>2013</year>
<fpage>4733</fpage>
<lpage>4738</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/IECON.2013.6699900">https://doi.org/10.1109/IECON.2013.6699900</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/IECON.2013.6699900</pub-id>
<conf-name>Are domestic load profiles stable over time? An attempt to identify target households for demand side management campaigns</conf-name>
<conf-loc>IECON</conf-loc>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref10">
<mixed-citation>Capozzoli, A., Piscitelli, M., Brandi, S., Grassi, D., &amp; Chicco, G. (2018). Automated load pattern learning and anomaly detection for enhancing energy management in smart buildings. <italic>Energy, 157</italic>, 336-352. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.energy.2018.05.127">https://doi.org/10.1016/j.energy.2018.05.127</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Capozzoli</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Piscitelli</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Brandi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Grassi</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chicco</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>Automated load pattern learning and anomaly detection for enhancing energy management in smart buildings</article-title>
<source>Energy</source>
<year>2018</year>
<volume>157</volume>
<fpage>336</fpage>
<lpage>352</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.energy.2018.05.127">https://doi.org/10.1016/j.energy.2018.05.127</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.energy.2018.05.127</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref11">
<mixed-citation>Chicco, G. (2012). Overview and performance assessment of the clustering methods for electrical load pattern grouping. <italic>Energy, 42</italic>(1), 68-80. DOI: 10.1016/j.energy.2011.12.031</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chicco</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>Overview and performance assessment of the clustering methods for electrical load pattern grouping</article-title>
<source>Energy</source>
<year>2012</year>
<volume>42</volume>
<issue>1</issue>
<fpage>68</fpage>
<lpage>80</lpage>
<pub-id pub-id-type="doi">10.1016/j.energy.2011.12.031</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref12">
<mixed-citation>Defays, D. (1977). An efficient algorithm for a complete link method. <italic>The Computer Journal, 20</italic>(4), 364-366. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/comjnl/20.4.364">https://doi.org/10.1093/comjnl/20.4.364</ext-link>.</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Defays</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>An efficient algorithm for a complete link method</article-title>
<source>The Computer Journal</source>
<year>1977</year>
<volume>20</volume>
<issue>4</issue>
<fpage>364</fpage>
<lpage>366</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/comjnl/20.4.364">https://doi.org/10.1093/comjnl/20.4.364</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1093/comjnl/20.4.364</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref13">
<mixed-citation>Dent, I., Aickelin, U., &amp; Rodden, T. (2011). Application of a clustering framework to UK domestic electricity data. <italic>Ukci</italic>, 161-166. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1307.1079">https://arxiv.org/abs/1307.1079</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dent</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Aickelin</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Rodden</surname>
<given-names>T.</given-names>
</name>
</person-group>
<article-title>Application of a clustering framework to UK domestic electricity data</article-title>
<source>Ukci</source>
<year>2011</year>
<fpage>161</fpage>
<lpage>166</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1307.1079">https://arxiv.org/abs/1307.1079</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref14">
<mixed-citation>Espinoza, M., Joye, C., Belmans, R., &amp; DeMoor, B. (2005). Short-Term load forecasting, profile identification, and customer segmentation: A methodology based on periodic time series. <italic>IEEE Transaction. Power Systems, 20</italic>(3), 1622-1630. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRS.2005.852123">https://doi.org/10.1109/TPWRS.2005.852123</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Espinoza</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Joye</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Belmans</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>DeMoor</surname>
<given-names>B.</given-names>
</name>
</person-group>
<article-title>Short-Term load forecasting, profile identification, and customer segmentation: A methodology based on periodic time series</article-title>
<source>IEEE Transaction Power Systems</source>
<year>2005</year>
<volume>20</volume>
<issue>3</issue>
<fpage>1622</fpage>
<lpage>1630</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRS.2005.852123">https://doi.org/10.1109/TPWRS.2005.852123</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/TPWRS.2005.852123</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref15">
<mixed-citation>European Commission. (2014). Benchmarking smart metering deployment in the EU-27 with a focus on electricity. <italic>Reports</italic>. Publications Office of the European Union <ext-link ext-link-type="uri" xlink:href="https://ses.jrc.ec.europa.eu/publications/reports/benchmarking-smart-metering-deployment-eu-27-focus-electricity">https://ses.jrc.ec.europa.eu/publications/reports/benchmarking-smart-metering-deployment-eu-27-focus-electricity</ext-link>
</mixed-citation>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<collab>European Commission</collab>
</person-group>
<article-title>Benchmarking smart metering deployment in the EU-27 with a focus on electricity</article-title>
<source>Reports</source>
<year>2014</year>
<publisher-name>Publications Office of the European Union</publisher-name>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://ses.jrc.ec.europa.eu/publications/reports/benchmarking-smart-metering-deployment-eu-27-focus-electricity">https://ses.jrc.ec.europa.eu/publications/reports/benchmarking-smart-metering-deployment-eu-27-focus-electricity</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref16">
<mixed-citation>Fenza, G., Gallo, M., &amp; Loia, V. (2019). Drift-aware methodology for anomaly detection in smart grid. <italic>IEEE Access, 7</italic>, 9645-9657. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/ACCESS.2019.2891315">https://doi.org/10.1109/ACCESS.2019.2891315</ext-link>
</mixed-citation>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Fenza</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gallo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Loia</surname>
<given-names>V.</given-names>
</name>
</person-group>
<article-title>Drift-aware methodology for anomaly detection in smart grid</article-title>
<source>IEEE Access</source>
<year>2019</year>
<volume>7</volume>
<fpage>9645</fpage>
<lpage>9657</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/ACCESS.2019.2891315">https://doi.org/10.1109/ACCESS.2019.2891315</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/ACCESS.2019.2891315</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref17">
<mixed-citation>Figueiredo, V., Rodrigues, F. Vale, Z., &amp; Gouveia, J. (2005). An electric energy consumer characterization framework based on data mining techniques. <italic>IEEE Transaction. Power Systems, 20</italic>(2), 596-602. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRS.2005.846234">https://doi.org/10.1109/TPWRS.2005.846234</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Figueiredo</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Rodrigues</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Vale</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gouveia</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>An electric energy consumer characterization framework based on data mining techniques</article-title>
<source>IEEE Transaction Power Systems</source>
<year>2005</year>
<volume>20</volume>
<issue>2</issue>
<fpage>596</fpage>
<lpage>602</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRS.2005.846234">https://doi.org/10.1109/TPWRS.2005.846234</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/TPWRS.2005.846234</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref18">
<mixed-citation>Flath, C., Nicolay, D., Conte, T., Van Dinther, C., &amp; Filipova-Neumann, L. (2012). Cluster analysis of smart metering data: An implementation in practice. <italic>Business &amp; Information Systems Engineering, 4</italic>, 31-39. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s12599-011-0201-5">https://doi.org/10.1007/s12599-011-0201-5</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Flath</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Nicolay</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Conte</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Van Dinther</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Filipova-Neumann</surname>
<given-names>L.</given-names>
</name>
</person-group>
<article-title>Cluster analysis of smart metering data: An implementation in practice</article-title>
<source>Business &amp; Information Systems Engineering</source>
<year>2012</year>
<volume>4</volume>
<fpage>31</fpage>
<lpage>39</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s12599-011-0201-5">https://doi.org/10.1007/s12599-011-0201-5</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1007/s12599-011-0201-5</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref19">
<mixed-citation>Funde, N., Dhabu, M., Paramasivam, A., &amp; Deshpande, P. (2019). Motif-based association rule mining and clustering technique for determining energy usage patterns for smart meter data. <italic>Sustainable Cities Society, 46</italic>(January), 101415. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.scs.2018.12.043">https://doi.org/10.1016/j.scs.2018.12.043</ext-link>
</mixed-citation>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Funde</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Dhabu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Paramasivam</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Deshpande</surname>
<given-names>P.</given-names>
</name>
</person-group>
<article-title>Motif-based association rule mining and clustering technique for determining energy usage patterns for smart meter data</article-title>
<source>Sustainable Cities Society</source>
<year>2019</year>
<volume>46</volume>
<issue>January</issue>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.scs.2018.12.043">https://doi.org/10.1016/j.scs.2018.12.043</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.scs.2018.12.043</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref20">
<mixed-citation>Giordano, V., Gangale, F., Fulli, G., &amp; Sánchez, M. (2011). <italic>Smart grids projects in Europe: Lessons learned and current developments</italic>. Federal Energy Regulatory Commission. <ext-link ext-link-type="uri" xlink:href="https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf">https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf</ext-link>
</mixed-citation>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Giordano</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Gangale</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Fulli</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sánchez</surname>
<given-names>M.</given-names>
</name>
</person-group>
<source>Smart grids projects in Europe: Lessons learned and current developments</source>
<year>2011</year>
<publisher-name>Federal Energy Regulatory Commission</publisher-name>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf">https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref21">
<mixed-citation>Hossain, J., Kabir, A., Rahman, M., Kabir, B., &amp; Islam, R. (2011). Determination of typical load profile of consumers using fuzzy c-means clustering algorithm. <italic>Int. J. Soft Comput. Eng</italic>., .(5), 169-173. <ext-link ext-link-type="uri" xlink:href="https://es.scribd.com/document/349605721/Determination-of-Typical-Load-Profile-of-Consumers-Using-Fuzzy-C-Means-Clustering-Algorithm">https://es.scribd.com/document/349605721/Determination-of-Typical-Load-Profile-of-Consumers-Using-Fuzzy-C-Means-Clustering-Algorithm</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hossain</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kabir</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rahman</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kabir</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Islam</surname>
<given-names>R.</given-names>
</name>
</person-group>
<article-title>Determination of typical load profile of consumers using fuzzy c-means clustering algorithm</article-title>
<source>Int J Soft Comput Eng</source>
<year>2011</year>
<volume>1</volume>
<issue>5</issue>
<fpage>169</fpage>
<lpage>173</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://es.scribd.com/document/349605721/Determination-of-Typical-Load-Profile-of-Consumers-Using-Fuzzy-C-Means-Clustering-Algorithm">https://es.scribd.com/document/349605721/Determination-of-Typical-Load-Profile-of-Consumers-Using-Fuzzy-C-Means-Clustering-Algorithm</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref22">
<mixed-citation>Hübner, M., &amp; Prüggler, N. (2011). Smart grids initiatives in Europe - Country snapshots and country fact sheets. <italic>JRC Reference Reports</italic>. Austrian Energy Agency. <ext-link ext-link-type="uri" xlink:href="https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf">https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf</ext-link>
</mixed-citation>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Hübner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Prüggler</surname>
<given-names>N.</given-names>
</name>
</person-group>
<article-title>Smart grids initiatives in Europe - Country snapshots and country fact sheets</article-title>
<source>JRC Reference Reports</source>
<year>2011</year>
<publisher-name>Austrian Energy Agency</publisher-name>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf">https://ses.jrc.ec.europa.eu/sites/ses/files/documents/smart_grid_projects_in_europe.pdf</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref23">
<mixed-citation>Iglesias, F., &amp; Kastner, W. (2013). Analysis of similarity measures in times series clustering for the discovery of building energy patterns. <italic>Energies</italic>, 6, 579-597. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/en6020579">https://doi.org/10.3390/en6020579</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Iglesias</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kastner</surname>
<given-names>W.</given-names>
</name>
</person-group>
<article-title>Analysis of similarity measures in times series clustering for the discovery of building energy patterns</article-title>
<source>Energies</source>
<year>2013</year>
<volume>6</volume>
<fpage>579</fpage>
<lpage>597</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/en6020579">https://doi.org/10.3390/en6020579</ext-link>
</comment>
<pub-id pub-id-type="doi">10.3390/en6020579</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref24">
<mixed-citation>Kohonen, T. (1982). Self-organized formation of topologically correct feature maps. <italic>Biol. Cybern</italic>., <italic>43</italic>(1), 59-69. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/BF00337288">https://doi.org/10.1007/BF00337288</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kohonen</surname>
<given-names>T.</given-names>
</name>
</person-group>
<article-title>Self-organized formation of topologically correct feature maps</article-title>
<source>Biol Cybern</source>
<year>1982</year>
<volume>43</volume>
<issue>1</issue>
<fpage>59</fpage>
<lpage>69</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/BF00337288">https://doi.org/10.1007/BF00337288</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1007/BF00337288</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref25">
<mixed-citation>Lai, C.-P., Chung, P.-C., &amp; Tseng, V. S. (2010). A novel two-level clustering method for time series data analysis. <italic>Expert Systems with Applications, 37</italic>(9), 6319-6326. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.eswa.2010.02.089">https://doi.org/10.1016/j.eswa.2010.02.089</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lai</surname>
<given-names>C.-P.</given-names>
</name>
<name>
<surname>Chung</surname>
<given-names>P.-C.</given-names>
</name>
<name>
<surname>Tseng</surname>
<given-names>V. S.</given-names>
</name>
</person-group>
<article-title>A novel two-level clustering method for time series data analysis</article-title>
<source>Expert Systems with Applications</source>
<year>2010</year>
<volume>37</volume>
<issue>9</issue>
<fpage>6319</fpage>
<lpage>6326</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.eswa.2010.02.089">https://doi.org/10.1016/j.eswa.2010.02.089</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.eswa.2010.02.089</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref26">
<mixed-citation>Lavin, A., &amp; Klabjan, D. (2014). Clustering time ‐ series energy data from smart meters. <italic>Energy Efficiency, 8</italic>(4), 1-9. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s12053-014-9316-0">https://doi.org/10.1007/s12053-014-9316-0</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lavin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Klabjan</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>Clustering time ‐ series energy data from smart meters</article-title>
<source>Energy Efficiency</source>
<year>2014</year>
<volume>8</volume>
<issue>4</issue>
<fpage>1</fpage>
<lpage>9</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s12053-014-9316-0">https://doi.org/10.1007/s12053-014-9316-0</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1007/s12053-014-9316-0</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref27">
<mixed-citation>Lee, T., Haben, S., &amp; Grindrod, P. (2014). <italic>Modelling the electricity consumption of small to medium enterprises</italic>. In The 18th European Conference on Mathematics for Industry Conference (pp. 1-7). Taormina, Italy: ECMI.</mixed-citation>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Haben</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Grindrod</surname>
<given-names>P.</given-names>
</name>
</person-group>
<source>The 18th European Conference on Mathematics for Industry Conference</source>
<year>2014</year>
<fpage>1</fpage>
<lpage>7</lpage>
<conf-name>Modelling the electricity consumption of small to medium enterprises</conf-name>
<conf-loc>Taormina, Italy: ECMI</conf-loc>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref28">
<mixed-citation>Losa, I., De Nigris, M., &amp; Van, T. (2013). Analysis of the on-going research and demonstration efforts on smart grids in Europe, 22nd International Conference on Electricity Distribution, June, 10-13. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1049/cp.2013.0958">https://doi.org/10.1049/cp.2013.0958</ext-link>
</mixed-citation>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Losa</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>De Nigris</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Van</surname>
<given-names>T.</given-names>
</name>
</person-group>
<source>22nd International Conference on Electricity Distribution</source>
<year>2013</year>
<month>06</month>
<fpage>10</fpage>
<lpage>13</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1049/cp.2013.0958">https://doi.org/10.1049/cp.2013.0958</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1049/cp.2013.0958</pub-id>
<conf-name>Analysis of the on-going research and demonstration efforts on smart grids in Europe</conf-name>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref29">
<mixed-citation>MacQueen, J. (1967). <italic>Some methods for classification and analysis of multivariate observations</italic>. In Proceedings of the 5th Berkeley Symposium Mathematical Statistics and Probability, 1, 281-297.</mixed-citation>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>MacQueen</surname>
<given-names>J.</given-names>
</name>
</person-group>
<source>Proceedings of the 5th Berkeley Symposium Mathematical Statistics and Probability</source>
<year>1967</year>
<volume>1</volume>
<fpage>281</fpage>
<lpage>297</lpage>
<conf-name>Some methods for classification and analysis of multivariate observations</conf-name>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref30">
<mixed-citation>McLoughlin, F., Duffy, A., &amp; Conlon, M. (2012). <italic>Analysing domestic electricity smart metering data using self organising maps</italic>. In CIRED 2012 Workshop: Integration of Renewables into the Distribution Grid, pp. 319-319. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1049/cp.2012.0865">https://doi.org/10.1049/cp.2012.0865</ext-link>
</mixed-citation>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>McLoughlin</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Duffy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Conlon</surname>
<given-names>M.</given-names>
</name>
</person-group>
<source>CIRED 2012 Workshop: Integration of Renewables into the Distribution Grid</source>
<year>2012</year>
<fpage>319</fpage>
<lpage>319</lpage>
<chapter-title>Analysing domestic electricity smart metering data using self organising maps</chapter-title>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1049/cp.2012.0865">https://doi.org/10.1049/cp.2012.0865</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1049/cp.2012.0865</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref31">
<mixed-citation>Mutanen, A., Ruska, M., Repo, S., &amp; Järventausta, P. (2011). Customer classification and load profiling method for distribution systems. <italic>IEEE Trans. Power Deliv</italic>., <italic>26</italic>, 1755-1763. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRD.2011.2142198">https://doi.org/10.1109/TPWRD.2011.2142198</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mutanen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ruska</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Repo</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Järventausta</surname>
<given-names>P.</given-names>
</name>
</person-group>
<article-title>Customer classification and load profiling method for distribution systems</article-title>
<source>IEEE Trans Power Deliv</source>
<year>2011</year>
<volume>26</volume>
<fpage>1755</fpage>
<lpage>1763</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/TPWRD.2011.2142198">https://doi.org/10.1109/TPWRD.2011.2142198</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/TPWRD.2011.2142198</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref32">
<mixed-citation>Nanopoulos, A., Alcock, R., &amp; Manolopoulos, Y. (2001). Feature-based classification of time-series data. <italic>International Journal of Computer Research</italic>, 10, 49-61. <ext-link ext-link-type="uri" xlink:href="http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.73.9555">http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.73.9555</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nanopoulos</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Alcock</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Manolopoulos</surname>
<given-names>Y.</given-names>
</name>
</person-group>
<article-title>Feature-based classification of time-series data</article-title>
<source>International Journal of Computer Research</source>
<year>2001</year>
<volume>10</volume>
<fpage>49</fpage>
<lpage>61</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.73.9555">http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.73.9555</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref33">
<mixed-citation>Oates, T., Firoiu, L., &amp; Cohen, P. (1999). <italic>Clustering time series with hidden Markov Models and dynamic time warping</italic>. In Proceedings of the IJCAI-99 Workshop on Neural, symbolic and reinforcement learning methods for sequence learning (pp. 17-21).</mixed-citation>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Oates</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Firoiu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cohen</surname>
<given-names>P.</given-names>
</name>
</person-group>
<source>Proceedings of the IJCAI-99 Workshop on Neural, symbolic and reinforcement learning methods for sequence learning</source>
<year>1999</year>
<fpage>17</fpage>
<lpage>21</lpage>
<conf-name>Clustering time series with hidden Markov Models and dynamic time warping</conf-name>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref34">
<mixed-citation>Räsänen, T., &amp; Kolehmainen, M. (2009). <italic>Feature-based clustering for electricity use time series data</italic>. In 9th International Conference, ICANNGA 2009, LNCS 5495, 401-412. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-642-04921-7_41">https://doi.org/10.1007/978-3-642-04921-7_41</ext-link>
</mixed-citation>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Räsänen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kolehmainen</surname>
<given-names>M.</given-names>
</name>
</person-group>
<source>9th International Conference</source>
<year>2009</year>
<fpage>401</fpage>
<lpage>412</lpage>
<comment>LNCS 5495</comment>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-642-04921-7_41">https://doi.org/10.1007/978-3-642-04921-7_41</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1007/978-3-642-04921-7_41</pub-id>
<conf-name>Feature-based clustering for electricity use time series data</conf-name>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref35">
<mixed-citation>Räsänen, T., Voukantsis, D., Niska, H., Karatzas, K., &amp; Kolehmainen, M. (2010). Data-based method for creating electricity use load profiles using large amount of customer-specific hourly measured electricity use data. <italic>Appl. Energy, 87</italic>, 3538-3545. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.apenergy.2010.05.015">https://doi.org/10.1016/j.apenergy.2010.05.015</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Räsänen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Voukantsis</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Niska</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Karatzas</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kolehmainen</surname>
<given-names>M.</given-names>
</name>
</person-group>
<article-title>Data-based method for creating electricity use load profiles using large amount of customer-specific hourly measured electricity use data</article-title>
<source>Appl Energy</source>
<year>2010</year>
<volume>87</volume>
<fpage>3538</fpage>
<lpage>3545</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.apenergy.2010.05.015">https://doi.org/10.1016/j.apenergy.2010.05.015</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.apenergy.2010.05.015</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref36">
<mixed-citation>Renner, S., &amp; Heinemann, C. (2011). <italic>European Smart Metering Landscape Report</italic>, 2(February), 168 p. <ext-link ext-link-type="uri" xlink:href="https://www.sintef.no/globalassets/project/smartregions/d2.1_european-smart-metering-landscape-report_final.pdf">https://www.sintef.no/globalassets/project/smartregions/d2.1_european-smart-metering-landscape-report_final.pdf</ext-link>
</mixed-citation>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Renner</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Heinemann</surname>
<given-names>C.</given-names>
</name>
</person-group>
<source>European Smart Metering Landscape Report</source>
<year>2011</year>
<volume>2</volume>
<issue>February</issue>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://www.sintef.no/globalassets/project/smartregions/d2.1_european-smart-metering-landscape-report_final.pdf">https://www.sintef.no/globalassets/project/smartregions/d2.1_european-smart-metering-landscape-report_final.pdf</ext-link>
</comment>
<size units="pages">168</size>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref37">
<mixed-citation>Shi, J., &amp; Malik, J. (2000). Normalized cuts and image segmentation. <italic>IEEE Trans. Pattern Anal. Mach. Intell</italic>., <italic>22</italic>(8), 888-905. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/34.868688">https://doi.org/10.1109/34.868688</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Malik</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>Normalized cuts and image segmentation</article-title>
<source>IEEE Trans Pattern Anal Mach Intell</source>
<year>2000</year>
<volume>22</volume>
<issue>8</issue>
<fpage>888</fpage>
<lpage>905</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/34.868688">https://doi.org/10.1109/34.868688</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1109/34.868688</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref38">
<mixed-citation>Wang, X., Smith, K., &amp; Hyndman, R. (2006). Characteristic-based clustering for time series data. <italic>Data Min. Knowl. Discov</italic>., 13, 335-364. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10618-005-0039-x">https://doi.org/10.1007/s10618-005-0039-x</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Hyndman</surname>
<given-names>R.</given-names>
</name>
</person-group>
<article-title>Characteristic-based clustering for time series data</article-title>
<source>Data Min Knowl Discov</source>
<year>2006</year>
<volume>13</volume>
<fpage>335</fpage>
<lpage>364</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10618-005-0039-x">https://doi.org/10.1007/s10618-005-0039-x</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1007/s10618-005-0039-x</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref39">
<mixed-citation>Wang, X., Smith, K., Hyndman, R., &amp; Alahakoon, D. (2004). A scalable method for time series clustering. Research of Monash University. Victoria, Australia: Monash University. <ext-link ext-link-type="uri" xlink:href="http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.155.207">http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.155.207</ext-link>
</mixed-citation>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Hyndman</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Alahakoon</surname>
<given-names>D.</given-names>
</name>
</person-group>
<source>Research of Monash University</source>
<year>2004</year>
<publisher-loc>Victoria, Australia</publisher-loc>
<publisher-name>Monash University</publisher-name>
<chapter-title>A scalable method for time series clustering. Research of Monash University</chapter-title>
<comment>
<ext-link ext-link-type="uri" xlink:href="http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.155.207">http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.155.207</ext-link>
</comment>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref40">
<mixed-citation>Warren Liao, T. (2007). A clustering procedure for exploratory mining of vector time series. <italic>Pattern Recognit</italic>., 40, 2550-2562. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.patcog.2007.01.005">https://doi.org/10.1016/j.patcog.2007.01.005</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Warren Liao</surname>
<given-names>T.</given-names>
</name>
</person-group>
<article-title>A clustering procedure for exploratory mining of vector time series</article-title>
<source>Pattern Recognit</source>
<year>2007</year>
<volume>40</volume>
<fpage>2550</fpage>
<lpage>2562</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.patcog.2007.01.005">https://doi.org/10.1016/j.patcog.2007.01.005</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.patcog.2007.01.005</pub-id>
</element-citation>
</ref>
<ref id="redalyc_20562876001_ref41">
<mixed-citation>Zhang, X., Liu, J., Du, Y., &amp; Lv, T. (2011). A novel clustering method on time series data. <italic>Expert Systems Applications, 38</italic>, 11891-11900. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.eswa.2011.03.081">https://doi.org/10.1016/j.eswa.2011.03.081</ext-link>
</mixed-citation>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lv</surname>
<given-names>T.</given-names>
</name>
</person-group>
<article-title>A novel clustering method on time series data</article-title>
<source>Expert Systems Applications</source>
<year>2011</year>
<volume>38</volume>
<fpage>11891</fpage>
<lpage>11900</lpage>
<comment>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.eswa.2011.03.081">https://doi.org/10.1016/j.eswa.2011.03.081</ext-link>
</comment>
<pub-id pub-id-type="doi">10.1016/j.eswa.2011.03.081</pub-id>
</element-citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn1" fn-type="other">
<label>*</label>
<p>Research paper</p>
</fn>
</fn-group>
</back>
</article>