[DATA][ENG.] 2. Why Do We Hash Data? Why Is Salt Necessary?
When working with data, you often encounter requirements like these:
- “This value must uniquely identify an entity, but the original value must not be visible.”
- “We only need to know whether different data belongs to the same user.”
- “Personal information must never be stored.”
In such cases, the most common solution is hashing.
However, in many situations, people simply think:
“Isn’t this just for security?”
and move on without fully understanding it.
In this article, we’ll explain why data hashing is necessary, what salt is, and when hashing should—and should not—be used.
What Is a Hash?
A hash is a function that converts an input value into a fixed-length output. A key property is that the original value can never be recovered from the hash.
Example:
- Input: user_id = "abc123"
- Output: e99a18c428cb38d5f260853678922e03
No matter how long or short the input is, the output length is always the same.
Its characteristics are clear:
- Same input → always the same output
- Even a tiny change in input → completely different output
- Impossible to restore the original value from the result (one-way)
This one-way property is the core of hashing.
Why Do We Hash Data?
1. Privacy Protection
This is the most common reason.
- Email addresses
- Phone numbers
- National ID numbers
- Customer IDs
These values are often needed for analysis, but storing the raw values is risky.
- email → hash(email)
- phone → hash(phone)
This allows you to:
- Identify whether records belong to the same user
- Avoid exposing actual personal information
In other words, hashing creates a state that is “identifiable but not recognizable.”
2. Unifying Internal Identifiers
In real-world systems, situations like this are common:
- user_id in Service A
- customer_no in Service B
- device_id in a logging system
When you want to unify different ID schemes into a single key, hashing provides a clean solution:
- hash(user_id)
- hash(device_id)
- hash(customer_no)
→ All mapped into the same hash space
3. Data Sharing / External Distribution
When data must be shared externally—for example:
- With partners
- For outsourced analysis
- For research purposes
Passing raw IDs is usually unacceptable.
Hashed data, however:
- Cannot directly identify individuals once exported
- Can still be matched internally via controlled mapping
- Significantly reduces legal and compliance risk
That’s why, in practice, people often say:
“Data that isn’t hashed can’t leave the company.”
4. Machine Learning / Analytical Stability
This is a surprisingly important reason.
Models often treat IDs as meaningless numbers:
- user_id = 1001
- user_id = 1002
The numeric magnitude itself has no semantic meaning, but models may mistakenly interpret it as meaningful.
By using hashes:
- Randomness is introduced
- Order and magnitude lose meaning
- The risk of models memorizing specific IDs is reduced
5. The Principle of Data Minimization
One core principle of good data design is:
Store only what you actually need.
Hashing fits this principle well.
- You don’t need to know who the person is
- You only need to know whether records belong to the same person
In such cases, storing the original data is excessive.
A hash retains exactly what is necessary—and nothing more.
Hash ≠ Encryption
There is a critical misconception that must be clarified.
|
Category |
Hash |
Encryption |
|---|---|---|
|
Reversible |
❌ No |
✅ Yes |
|
Purpose |
Identification / Protection |
Confidentiality |
|
Key Required |
❌ No |
✅ Yes |
|
Original Recovery |
❌ Impossible |
✅ Possible |
When You Should
Not
Use Hashing
In the following cases, hashing can actually cause problems:
- When you need to recover the original value later
- When personal data must be provided due to legal requests
- When the input domain is very small (e.g., gender, Yes/No, country codes)
If the value range is small, brute-force attacks become feasible.
Even though hashes cannot be reversed directly, an attacker can repeatedly hash all possible inputs until a matching hash is found.
So, When Is Salt Necessary?
A key concept often mentioned alongside hashing is salt.
Because:
- Same value → same hash
this property introduces vulnerabilities:
- Brute-force attacks
- Dictionary attacks (rainbow tables)
To mitigate this, we use salt:
hash(value + salt)In other words, we “add a bit of salt” to the input.
In practice:
- The salt is usually a fixed value
- It is kept secret by the team that generates the hashed data
General guidelines:
- Externally shared data → Salt is mandatory
- Internal analytical IDs → Optional, depending on the use case
Commonly Used Hash Algorithms
For security-related hashing, the SHA family is widely used.
Key characteristics:
- Extremely low collision probability
- Highly sensitive to input changes
- Guaranteed one-way behavior
SHA-256 / SHA-512 (De Facto Standards)
Among SHA algorithms, SHA-256 and SHA-512 are the most commonly used.
- Differentiated by output length: 256-bit / 512-bit
- Widely adopted industry standards
- Suitable for data identification, integrity checks, and ID hashing
Example:
SHA-256("hello") → 2cf24dba5fb0...✔ Typical use cases:
- Hashing personal data
- File integrity verification
- Internal identifier generation