Dalarna University's logo and link to the university's website

du.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • chicago-author-date
  • chicago-note-bibliography
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
En jämförelse av hanteringen för saknade data inom EDA
Dalarna University, School of Information and Engineering.
Dalarna University, School of Information and Engineering.
2024 (Swedish)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesisAlternative title
How the approach to handling missing values affects the result for Exploratory data analysis (English)
Abstract [sv]

This thesis aims to explore different approaches to handling missing values for Exploratory Data Analysis. This is done to illustrate the importance of picking a suitable approach for any giving dataset. The motivation for this stems from the importance of showing an accurate representation of data. Small deviations in the result can have major implications depending on the subject regarding the dataset. To answer this question this thesis shows the process of an EDA, and it will document the differences in results between these approaches. Data normalization (if applicable) and the handling of empty datapoints are explored and we found that the approach can have a significant impact on the result. How significant the impact is varies depending on the number of missing values and approach. With a larger number of missing values, the impact is very significant compared to when there are fewer missing values since more data changes. We can conclude that the approach does matter but it is equally important to consider the number of missing values in the dataset and where these values are located.

The subject for the EDA in question is Co2 emission rates per country and the dataset used contains 61,425 rows and 12 columns. The size of the dataset is important since too small a sample size could potentially show inaccurate results if that sample contains outliers. The analysis used the programming language Python to conduct all handling and visualization of data and Jupyter Notebook for structuring of code.  

Place, publisher, year, edition, pages
2024.
Keywords [sv]
Datahantering, Exploratory Data Analysis, Missande data
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:du-49051OAI: oai:DiVA.org:du-49051DiVA, id: diva2:1883432
Subject / course
Microdata Analysis
Available from: 2024-07-10 Created: 2024-07-10 Last updated: 2025-10-09

Open Access in DiVA

fulltext(820 kB)195 downloads
File information
File name FULLTEXT01.pdfFile size 820 kBChecksum SHA-512
2de10bddaf39595afbe04a72a091fea987f941645c23c1a3c27528d83b6a41d235492d8a96edacaa98870f25a0d69619bfdaa9005556b0f39f691bb8d53cee6f
Type fulltextMimetype application/pdf

By organisation
School of Information and Engineering
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 197 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 355 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • chicago-author-date
  • chicago-note-bibliography
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf