Table of Contents
Machine learning depends on data that matches the question you want a model to answer. Before choosing an algorithm, define the outcome, collect relevant examples, check their quality, and set aside data for evaluation. There is no universal percentage of project time spent on collection or cleaning.
What counts as machine learning data?
A dataset contains observations, often represented as rows. Its columns may hold features used for prediction and a target you want the model to learn. To estimate a home's sale price, for example, floor area and location can be features while the observed sale price is the target. Data can also be images, audio, text, or sequences rather than a simple spreadsheet.
| Data form | Example in a housing dataset | Typical preparation |
|---|---|---|
| Numeric | Floor area, sale price | Check units, missing values, and outliers |
| Categorical | Neighborhood or property type | Standardize labels and encode categories |
| Text | Listing description | Remove irrelevant text and protect personal details |
| Date or time | Sale date | Use a consistent format and avoid future information |
Quantitative variables are measured or counted numerically; qualitative or categorical variables describe classes or properties. A numeric code used to label a neighborhood is still a category, not a meaningful quantity.
A small example: floor area and price
The paired columns below are a toy dataset for illustration. Each position in the price row corresponds to the same position in the size row. Units and provenance were not specified, so these values illustrate layout rather than a real pricing relationship.
| Price | 7 | 8 | 8 | 9 | 9 | 9 | 10 | 11 | 14 | 14 | 15 |
| Size | 50 | 60 | 70 | 80 | 90 | 100 | 110 | 120 | 130 | 140 | 150 |
Before training, give each column a name and unit, verify that price and size refer to the same property, and look for missing or implausible values. Then split observations so the model is tested on examples it did not train on. A tiny, unverified sample like this cannot support a reliable valuation.
Population, census, and sample
The population is the full group you want to understand. A census measures every member; a sample measures some of them. In a simple random sample, each member has an equal chance of selection. Sampling introduces uncertainty, but it is not automatically inaccurate: a well-designed sample can support useful estimates, while a full census can still contain measurement errors.
Selection bias occurs when the sampled group systematically differs from the population in a way that matters. For example, a survey shared only online may miss people with limited internet access. In machine learning, training only on one neighborhood may make a housing model weak elsewhere. See the statistics and sampling course for deeper practice.
A practical data-preparation checklist
- Define the prediction task, target, unit of observation, and intended population.
- Record where the data came from and whether its collection and use are permitted.
- Inspect missing values, duplicates, inconsistent units, and labels.
- Keep training and evaluation data separate; for time-based predictions, avoid letting later observations leak into earlier ones.
- Measure performance on relevant groups and document gaps before using the model for decisions.
Healthcare, finance, and business data all require this same discipline, plus safeguards appropriate to their sensitivity. Large-scale data processing and data mining are related methods for finding patterns, but a larger dataset does not fix poor labels or biased collection. Continue with TipsMake's machine learning fundamentals course to build and evaluate a baseline model.
Reader Comments 0
Sign in with email or Google to join the discussion.