AP® Cybersecurity review sheet from Aim for Five (aimforfive.com/cybersecurity/units/5/5-6)
Unit 5 · Topic 5.6
5.6 Detecting Attacks on Data and Applications
The last line of defense is noticing when someone touches, changes or steals data. This topic covers accounting logs, honeypot files, data loss prevention and hash checks, how to weigh their cost, speed and blind spots, and how to spot SQL injection, XSS, buffer overflow and directory traversal in web logs.
Key terms
- accounting
- honeypot
- file hash check
- data loss prevention (DLP)
- web log indicators
Ways to detect attacks on data
Defenders have four main ways to notice someone touching data they shouldn't:
- Accounting: recording and monitoring what users do, including when data is accessed and by whom. Red flags in these logs: opening files that are rarely used, activity outside a user's normal pattern (time of day, location, device), and attempts to copy or delete sensitive files.
- Honeypots: files that look valuable (fake card numbers, PII or passwords) but hold only fake data. No one has a legitimate reason to open one, so any access triggers an alert.
- Hash checks: a cryptographic hash is repeatable, so hashing an unchanged file always gives the same digest. Record a file's hash, hash it again later and compare. A different hash means the file was altered. To get a file's SHA-256 hash: on Linux (Bash)
sha256sum testfile; in Windows PowerShellGet-FileHash testfile -Algorithm SHA256; on a Mac (zsh)shasum -a 256 testfile. - Data loss prevention (DLP): services, often bought from a third party, that monitor how data is accessed, used and sent across the organization to catch suspicious activity.
Choosing and judging detective controls
Watch more sensitive or critical data more closely, since adversaries target it most. Data classified as private, educational, health or financial often comes with legal or regulatory monitoring requirements.
| Control | Cost | When it detects | Blind spot |
|---|---|---|---|
| Honeypot | Low | Almost instantly, during the attack | Misses adversaries who never touch it |
| Hash check | Low | After the change has happened | Can't see data that was read or stolen without being changed |
| DLP service | High | Often during the attack, since many DLP tools alert in real time | Cost is the main drawback; coverage is strong |
| Automated real-time log analysis | Varies | During the attack | Needs good rules or models |
| Manual or after-the-fact log review | Low tools, high time | After the attack | Too slow alone; needs automation |
Reading web logs for application attacks
- SQL injection: quote characters (' or "), always-true conditions like OR 1=1, a double dash -- (starts a comment in SQL, which cuts off the rest of the real command), and SQL keywords like SELECT, WHERE, FROM or IN (usually written in capitals, though SQL accepts any case).
- Cross-site scripting: tags in user input, especially <script> … </script>.
- Buffer overflow: unusually long URLs, cookies, query strings or total request sizes.
- Directory traversal: GET requests whose paths contain runs of ../.
Layering the controls
These controls work best together. Picture a school district's HR file share. A honeypot file named 2026_staff_salaries_FINAL.xlsx, full of fake numbers, alerts the moment anyone opens it. Hashes of the payroll configuration files are recorded each night, so any change shows up the next morning. Accounting logs record who opened, copied or deleted files, so investigators can trace what an intruder touched. A DLP service, if the budget allows, watches for large amounts of data leaving the network. Each control covers a gap the others leave.
Worked examples
Try each one yourself first, then open the solution.
- Example 1
Has the file changed?
On Monday, an admin recorded the SHA-256 hash of payroll_rules.txt: b26148cdf027872fe36f215895558c5f67793b95edecf827496e214b97fabb36. On Friday, sha256sum payroll_rules.txt prints 50d7921d54717e474db005f2a14ddcdee52f142ae74457b25c0771513c0332f7. The file looks the same when opened. What can the admin conclude, and what can't they conclude?
Show the solutionHide the solution
- Step 1: Hashes are repeatable: an unchanged file always gives the same hash.
- Step 2: The two hashes differ, so the file was altered between Monday and Friday, even if the change is tiny. (Here, the only change was one added period.)
- Step 3: The hash can't say who changed it or why. Check the accounting logs for who accessed the file, and compare it to a backup to find the change.
- Step 4: Also note the blind spot: if someone had only copied the file, the hash would be unchanged.
Answer: The file was altered, because its hash changed. The hash doesn't show who changed it, what changed or whether anyone copied it, so the admin should review the access logs and compare the file to a backup.
- Example 2
Finding attacks in a web log
A web log from shop.example.com shows: Line 1: 198.51.100.14 GET /search?q=hiking+boots 200. Line 2: 203.0.113.9 GET /item?id=42' OR 1=1 -- 200. Line 3: 203.0.113.9 POST /reviews with comment text <script>(code removed)</script> 200. Line 4: 192.0.2.77 GET /images/../../../../etc/passwd 403. Line 5: 192.0.2.80 GET /login?user=AAAA… (a 9,000-character username) 500. Identify the attack on each suspicious line and cite the evidence.
Show the solutionHide the solution
- Step 1: Line 1: an ordinary search with normal words. No indicator.
- Step 2: Line 2: a quote, the always-true condition OR 1=1 and a -- comment marker inside the id parameter: SQL injection.
- Step 3: Line 3: a <script> tag submitted in a comment that will be shown to other visitors: stored XSS.
- Step 4: Line 4: a run of ../ sequences aimed at /etc/passwd: directory traversal (the 403 means the server refused it).
- Step 5: Line 5: a 9,000-character username, far longer than any real one, followed by a server error (500): an attempted buffer overflow.
- Step 6: Note that 203.0.113.9 appears on two attack lines, so it's worth blocking and investigating.
Answer: Line 2: SQL injection (quote, OR 1=1, --). Line 3: XSS (<script> tag in a comment). Line 4: directory traversal (../ toward /etc/passwd). Line 5: buffer overflow attempt (9,000-character input, server error). Line 1 is normal.
Common mistakes
- Thinking a hash check detects data theft. It only shows whether data changed; copying a file leaves its hash the same.
- Expecting honeypots to catch every intruder. They only alert when someone touches them.
- Flagging any request with a capital letter or a dash as SQL injection. Look for the combination: quotes, always-true conditions, -- and SQL keywords in input fields.
- Saying a 403 or 404 means there was no attack. The request is still evidence of an attempt, even if the server blocked it.
On the exam
- The free-response question gives you logs and file listings. Name each attack and cite the exact line number, IP address and the characters that give it away.
- When comparing detective controls, use cost, speed (during versus after the attack) and blind spots, like hashes missing theft and honeypots missing adversaries who avoid them.
Connected topics
Videos
Check yourself: 5.6 Detecting Attacks on Data and Applications
4 questions on 5.6 Detecting Attacks on Data and Applications. Pick an answer to see if you got it, and why.
Access log from an invented restaurant's web server; the script on line 3 has been shortened
Which line shows signs of an SQL injection attempt?
Which pair correctly matches lines 3 and 4 with the attack each suggests?
What does line 5 most likely indicate?
A company places a honeypot file named passwords_backup.xlsx on its file server. Which attack is the honeypot most likely to MISS?
0 of 4 answered