A content team with thousands of pages cannot refresh all of them at once. This paper builds a scoring system that ranks pages by their unrealized click potential given current search visibility. Using a 15,310-page slice of production search data (14 clients, 60-day window ending August 2025), a logistic regression with six pre-window features achieves precision@10 of 1.00 at identifying at-risk pages, compared to 0.20 for a weighted heuristic rule (base rate: 0.50). The model uses only data that exists before the label window, validated with GroupKFold by client to prevent memorization. The result is a ranked action playbook that tells a content team which pages to refresh first, with reason codes attached to every pick.
Consider a content team managing 25,000 search-facing pages. Some pages get impressions but no clicks. Others get both impressions and clicks. The team only has budget to refresh a few hundred pages per quarter. Which ones do they prioritize?
The obvious answer is "the ones with low click-through rate." But the problem is that CTR alone is misleading. A page with 10 impressions and 0 clicks is not the same as a page with 10,000 impressions and 50 clicks. The first page has no data and the second has a real problem.
This paper frames the question differently: which pages have the greatest unrealized click opportunity after accounting for their current search visibility? A page that gets 3,000 impressions but converts at 0.01% has more upside than a page that gets 50 impressions and converts at 2%. The first page is already being seen. It's just not being clicked.
The system scores every page, ranks them, and attaches a reason code to each pick. The output is not a prediction of future CTR. It is a prioritized list: start here, not there.
The FlyRank internship warehouse, a public snapshot hosted on HuggingFace at FlyRank/internship-warehouse. Three tables are used:
dim_content: page-level attributes (content type, main intent keyword search volume)fact_daily: daily search performance per page (impressions, clicks, average position)fact_query_90d: 90-day query-level aggregatesPages with at least 70 impressions in the prior 30 days, over a 60-day window ending 2025-08-31. This gives 15,310 content items across 14 clients. The 70-impression floor ensures enough data to compute a meaningful CTR. Clients with fewer than 100 pages in the slice are also excluded.
ctr_label = 100 x (clicks in last 30 days) / (impressions in last 30 days). A page is "at risk" if its CTR falls below the slice median. Base rate: 0.50 (balanced by construction).
Any column that references clicks or impressions within the label window is excluded to prevent leakage. This includes clicks_30d, impressions_last30d, and avg_position_30d. trend_direction and trend_pct are label-derived and also excluded. content_hash_id and client_hash_id are used only for grouping in GroupKFold, never as features. Features are strictly pre-window or static.
Six columns, all available before the label window:
| Feature | Type | Description |
|---|---|---|
search_volume | Numeric | Keyword demand estimate from dim_content |
impressions_prev30d | Numeric | Prior-30-day GSC impressions (days -60 to -31) |
has_search_volume | Binary | 1 if search_volume is missing, 0 otherwise |
has_main_intent | Binary | 1 if main_intent is missing, 0 otherwise |
main_intent | Categorical | informational, transactional, commercial, navigational |
content_type | Categorical | keyword article, feedly article |
Missing keyword volume is not imputed. The missingness itself is the signal. Pages with no keyword data tend to have different risk profiles. The categorical features are one-hot encoded, giving 10 model features total.
A weighted score that approximates what a human editor would do by sorting on impressions and keyword volume:
baseline = 0.50 x visibility_percentile + 0.45 x demand_percentile + 0.05 x transactional_flag
Missing keyword volume maps to zero demand. The rule is NaN-safe.
Logistic regression with max_iter=1000, random_state=42. No regularization tuning. The point is to test whether a simple linear model beats the rule, not to find the best possible model.
GroupKFold by client (5 folds). No client appears in both training and test sets in any fold. Fold sizes range from 907 to 7,735 test pages. This prevents the model from memorizing client-specific patterns.
Four demonstrations confirmed that features touching the label window produce inflated AUC. When clicks_30d is included, AUC reaches 0.854. With same-window impressions, 0.724. The honest feature set, restricted to pre-window data, achieves AUC 0.557. The gap between random-split and grouped-split AUC is the memorization cost.
The logistic regression outperforms the baseline rule at every cutoff on pages from clients the model never saw during training.
| Cutoff | Model | Rule | Base rate |
|---|---|---|---|
| p@10 | 1.00 | 0.20 | 0.50 |
| p@20 | 1.00 | 0.40 | 0.50 |
| p@50 | 0.94 | 0.58 | 0.50 |
| p@100 | 0.92 | 0.59 | 0.50 |
At p@10, the model places all 10 at-risk pages in the top 10. The rule places 2. The gap is largest at the top of the list, where a content team would actually start working.
| Fold | Train | Test | Model AUC | Baseline AUC |
|---|---|---|---|---|
| 1 | 7,575 | 7,735 | 0.586 | 0.513 |
| 2 | 12,356 | 2,954 | 0.554 | 0.513 |
| 3 | 12,613 | 2,697 | 0.591 | 0.513 |
| 4 | 14,293 | 1,017 | 0.687 | 0.513 |
| 5 | 14,403 | 907 | 0.562 | 0.513 |
| Mean | 0.596 | 0.513 |
The model beats the baseline on all five folds. Fold 4 (1,017 test pages) shows the highest AUC at 0.687, likely because the held-out client in that fold has a more distinct pattern.
Permutation importance on the full logistic regression shows that two features carry most of the signal:
| Feature | AUC drop |
|---|---|
impressions_prev30d | 0.039 |
search_volume | 0.030 |
main_intent_transactional | 0.020 |
has_search_volume | 0.005 |
The categorical features (intent, content type) add less than 0.02 each. Two features do the work: how much visibility the page already has, and how much demand the keyword table reports.
At the 90th percentile threshold (top 10% of model scores):
Random Forest added no value: p@10=0.62, p@20=0.61, p@50=0.61, essentially matching the rule. The extra nonlinear capacity earns no gain on near-monotone, low-signal features. This is a valid negative result. The relationship between features and label is approximately linear, and a more complex model does not help.
This work cannot claim the following:
The action playbook assigns each page a suggested action and reason codes. The top of the queue, ranked by model score:
Reason codes attached to every pick: recent_search_exposure (impressions >= 200), meaningful_demand (search_volume >= 1,000), transactional_priority.
The model should be re-evaluated when:
All code is in the GitHub repository under work/notebooks/capstone.ipynb. The notebook runs top to bottom with no errors. To reproduce:
pip install pandas scikit-learn matplotlib jupyterjupyter notebook work/notebooks/w07_action_playbook.ipynb to generate the feature vector work/outputs/action_playbook_queue.csvjupyter notebook work/notebooks/capstone.ipynbKey parameters: random_state=42, max_iter=1000, 5-fold GroupKFold by client. The data is loaded from work/outputs/action_playbook_queue.csv, a pre-built feature vector derived from the FlyRank warehouse.
Data provided by FlyRank AI, a public snapshot hosted on HuggingFace at FlyRank/internship-warehouse. The warehouse contains anonymized search performance data across 104 clients and 519,606 content items. This paper uses a 15,310-item slice with a 60-day window ending August 2025.
Built during the ML Engineering Internship, July - August 2026.
Library Versions::
All seeds set to 42 for reproducibility and CV=5.