Research · evaluated 9 October 2026
Evaluating URL-only phishing detection, honestly.
We trained models that look only at a URL’s text, tested them on websites they had never seen, and then tested them again on a different dataset. This page reports what held up and what did not.
Summary
- A shortcut, found and removed. In one dataset every legitimate URL was a bare homepage. A no-ML rule exploiting that scores F1 0.996 there. We excluded those signals from the main model.
- Better than the current demo on familiar data. On websites held out from the same dataset, the hostname model reaches ROC-AUC 0.919 against 0.848 for the model in today’s demo.
- Generalisation is the open problem. On a separate dataset the advantage shrinks to 0.741 against 0.717, with overlapping uncertainty, and a threshold tuned for 1% false positives produced 6.2%.
Data
Both datasets are published under CC BY 4.0, which permits this use with attribution. We use only each dataset’s URL text and label; supplied features derived from page content are ignored, because a URL-only detector cannot see them.
| Dataset | URLs | Phishing / legitimate | Registrable domains | Collected |
|---|---|---|---|---|
| PhiUSIIL (Prasad & Chandra, UCI #967) | 235,349 | 100,499 / 134,850 | 197,297 | Not stated (published 2024) |
| Web page phishing detection (Hannousse & Yahiouche, Mendeley v3) | 11,424 | 5,709 / 5,715 | 7,261 | May 2020 |
Cleaning removed URLs that the app’s own validator rejects, exact duplicates, and URLs with conflicting labels. Neither publisher documents where its URLs came from.
Method
- Split by owner, not by row. URLs are grouped by registrable domain using the Public Suffix List, then assigned 70/15/15 to training, validation and test. No domain appears in two splits.
- Choose on validation only. Model family, settings and decision thresholds were selected on the validation split. Test sets were scored once, afterwards.
- Test across datasets. Models trained on one dataset were also tested on the other, after removing every domain seen in training or validation.
- Report uncertainty honestly. Ranges are 95% intervals from a bootstrap that resamples whole domains, because a few domains contribute many URLs.
Three model families were compared: logistic regression on hand-designed lexical features (the baseline), gradient-boosted trees on the same features, and logistic regression on character 3–5-grams. The character n-gram model won on validation in both experiments.
Finding: a dataset shortcut
Before training anything, we profiled the URLs. PhiUSIIL’s legitimate class turned out to be built entirely from homepages:
| Property | Legitimate | Phishing |
|---|---|---|
| Uses https:// | 100% | 49% |
| Host starts with www. | 100% | 41% |
| Has a path or query | 0% | 28% |
A rule with no machine learning—legitimate if https, www and no path—reaches F1 0.996 on PhiUSIIL’s held-out domains and only 0.754 on the Hannousse dataset, where it flags 84% of legitimate URLs. A model that sees those properties would learn the dataset, not phishing. Our main model therefore sees only the hostname, with a leading www. removed.
Results
Positive class is phishing. “Recall at 1% FPR” uses the threshold chosen on validation to allow 1% false positives; the false-positive rate it actually produced on each test set is shown next to it.
Hostname model, trained on PhiUSIIL
| Test set | Model | ROC-AUC (95% interval) | PR-AUC | Recall at 1% FPR | Actual FPR |
|---|---|---|---|---|---|
| PhiUSIIL, held-out domains35,816 URLs · 29,522 domains | Hostname candidate | 0.919 (0.902–0.935) | 0.925 | 0.662 | 0.9% |
| Lexical baseline | 0.781 (0.689–0.849) | 0.827 | 0.469 | 1.1% | |
| Current demo model | 0.848 (0.809–0.885) | 0.840 | 0.530 | 0.7% | |
| Hannousse, different dataset9,625 URLs · 6,576 domains | Hostname candidate | 0.741 (0.711–0.772) | 0.769 | 0.356 | 6.2% |
| Lexical baseline | 0.664 (0.618–0.709) | 0.734 | 0.384 | 10.7% | |
| Current demo model | 0.717 (0.682–0.751) | 0.700 | 0.724 | 41.4% |
Full-URL model, trained on Hannousse
The Hannousse dataset has realistic legitimate URLs (with paths, both http and https), so here the model may read the whole URL. Its test set is small, so the intervals are wide.
| Test set | Model | ROC-AUC (95% interval) | PR-AUC | Recall at 1% FPR | Actual FPR |
|---|---|---|---|---|---|
| Hannousse, held-out domains1,969 URLs · 1,072 domains | Full-URL candidate | 0.967 (0.940–0.985) | 0.977 | 0.760 | 1.2% |
| Lexical baseline | 0.941 (0.904–0.971) | 0.957 | 0.574 | 1.2% | |
| Current demo model | 0.776 (0.651–0.858) | 0.822 | 0.000* | 0.0% |
* The current demo model gives its maximum score to more than 1% of legitimate validation URLs, so no threshold reaches 1% false positives and it flags nothing at that operating point.
The current demo, as it actually behaves
At the app’s own default cut-off (a model score of 0.5), the demo model catches 10.6% of phishing URLs on held-out PhiUSIIL domains, and 38.2% on the Hannousse dataset while flagging 13.4% of legitimate URLs. This is why the demo is labelled a demonstration.
Limitations
- Distribution shift. PhiUSIIL’s legitimate URLs are popular homepages; real-world legitimate links are far more varied, so false-positive rates in practice are likely higher than on PhiUSIIL.
- A few domains dominate. One hosting domain contributes 1,550 PhiUSIIL test URLs; one site is 18% of the Hannousse test set. That is why intervals are wide.
- No time dimension. Neither dataset has timestamps, so drift over time is unmeasured, and phishing infrastructure changes quickly.
- Text only. A URL-only model cannot recognise a compromised legitimate site, content on reputable hosting platforms, or a convincing new domain.
- Unknown provenance. Neither dataset documents its URL sources, so label noise cannot be quantified.
What comes next
The next step is improving cross-dataset behaviour, not the in-dataset score: training on both datasets together, weighting evaluation per domain, setting thresholds for each deployment context, and adding a licensed phishing source with timestamps to measure drift. Until then, the candidates stay experimental and the public demo keeps its warning labels.
Try the URL dissector · Open the research demo (opens in a new tab)
Data credits
Prasad, A. & Chandra, S. (2024). PhiUSIIL Phishing URL (Website) [Dataset]. UCI Machine Learning Repository. doi:10.1016/j.cose.2023.103545. CC BY 4.0.
Hannousse, A. & Yahiouche, S. (2021). Web page phishing detection (Version 3) [Dataset]. Mendeley Data. doi:10.17632/c2gw7fy2j4.3. CC BY 4.0.
Public Suffix List (Mozilla Public License 2.0), used for domain grouping and in the URL dissector.