Which Pages Have the Greatest Unrealized Click Opportunity?

ML Engineering Internship Capstone · July - August 2026 · Repo: ML_engineering_internship_project

Author: Emmanuel O. Ajala · GitHub · LinkedIn · Portfolio

A content team with thousands of pages cannot refresh all of them at once. This paper builds a scoring system that ranks pages by their unrealized click potential given current search visibility. Using a 15,310-page slice of production search data (14 clients, 60-day window ending August 2025), a logistic regression with six pre-window features achieves precision@10 of 1.00 at identifying at-risk pages, compared to 0.20 for a weighted heuristic rule (base rate: 0.50). The model uses only data that exists before the label window, validated with GroupKFold by client to prevent memorization. The result is a ranked action playbook that tells a content team which pages to refresh first, with reason codes attached to every pick.

1. Introduction

Consider a content team managing 25,000 search-facing pages. Some pages get impressions but no clicks. Others get both impressions and clicks. The team only has budget to refresh a few hundred pages per quarter. Which ones do they prioritize?

The obvious answer is "the ones with low click-through rate." But the problem is that CTR alone is misleading. A page with 10 impressions and 0 clicks is not the same as a page with 10,000 impressions and 50 clicks. The first page has no data and the second has a real problem.

This paper frames the question differently: which pages have the greatest unrealized click opportunity after accounting for their current search visibility? A page that gets 3,000 impressions but converts at 0.01% has more upside than a page that gets 50 impressions and converts at 2%. The first page is already being seen. It's just not being clicked.

The system scores every page, ranks them, and attaches a reason code to each pick. The output is not a prediction of future CTR. It is a prioritized list: start here, not there.

2. Data

Source

The FlyRank internship warehouse, a public snapshot hosted on HuggingFace at FlyRank/internship-warehouse. Three tables are used:

Slice

Pages with at least 70 impressions in the prior 30 days, over a 60-day window ending 2025-08-31. This gives 15,310 content items across 14 clients. The 70-impression floor ensures enough data to compute a meaningful CTR. Clients with fewer than 100 pages in the slice are also excluded.

Label

ctr_label = 100 x (clicks in last 30 days) / (impressions in last 30 days). A page is "at risk" if its CTR falls below the slice median. Base rate: 0.50 (balanced by construction).

Exclusions

Any column that references clicks or impressions within the label window is excluded to prevent leakage. This includes clicks_30d, impressions_last30d, and avg_position_30d. trend_direction and trend_pct are label-derived and also excluded. content_hash_id and client_hash_id are used only for grouping in GroupKFold, never as features. Features are strictly pre-window or static.

3. Methodology

Features

Six columns, all available before the label window:

FeatureTypeDescription
search_volumeNumericKeyword demand estimate from dim_content
impressions_prev30dNumericPrior-30-day GSC impressions (days -60 to -31)
has_search_volumeBinary1 if search_volume is missing, 0 otherwise
has_main_intentBinary1 if main_intent is missing, 0 otherwise
main_intentCategoricalinformational, transactional, commercial, navigational
content_typeCategoricalkeyword article, feedly article

Missing keyword volume is not imputed. The missingness itself is the signal. Pages with no keyword data tend to have different risk profiles. The categorical features are one-hot encoded, giving 10 model features total.

Baseline rule

A weighted score that approximates what a human editor would do by sorting on impressions and keyword volume:

baseline = 0.50 x visibility_percentile + 0.45 x demand_percentile + 0.05 x transactional_flag

Missing keyword volume maps to zero demand. The rule is NaN-safe.

Model

Logistic regression with max_iter=1000, random_state=42. No regularization tuning. The point is to test whether a simple linear model beats the rule, not to find the best possible model.

Validation

GroupKFold by client (5 folds). No client appears in both training and test sets in any fold. Fold sizes range from 907 to 7,735 test pages. This prevents the model from memorizing client-specific patterns.

Leakage checks

Four demonstrations confirmed that features touching the label window produce inflated AUC. When clicks_30d is included, AUC reaches 0.854. With same-window impressions, 0.724. The honest feature set, restricted to pre-window data, achieves AUC 0.557. The gap between random-split and grouped-split AUC is the memorization cost.

4. Results

Model vs baseline

The logistic regression outperforms the baseline rule at every cutoff on pages from clients the model never saw during training.

CutoffModelRuleBase rate
p@101.000.200.50
p@201.000.400.50
p@500.940.580.50
p@1000.920.590.50

At p@10, the model places all 10 at-risk pages in the top 10. The rule places 2. The gap is largest at the top of the list, where a content team would actually start working.

Precision at K chart
Figure 1. Model vs rule precision at K. The model maintains high precision even at p@100 (0.92), while the rule barely beats the base rate.

AUC across folds

FoldTrainTestModel AUCBaseline AUC
17,5757,7350.5860.513
212,3562,9540.5540.513
312,6132,6970.5910.513
414,2931,0170.6870.513
514,4039070.5620.513
Mean0.5960.513

The model beats the baseline on all five folds. Fold 4 (1,017 test pages) shows the highest AUC at 0.687, likely because the held-out client in that fold has a more distinct pattern.

Feature importance

Permutation importance on the full logistic regression shows that two features carry most of the signal:

FeatureAUC drop
impressions_prev30d0.039
search_volume0.030
main_intent_transactional0.020
has_search_volume0.005
Feature importance chart
Figure 2. Permutation importance (AUC drop when feature is shuffled). Two numeric features do most of the work.

The categorical features (intent, content type) add less than 0.02 each. Two features do the work: how much visibility the page already has, and how much demand the keyword table reports.

Error patterns

At the 90th percentile threshold (top 10% of model scores):

Surprises and Negative Results

Random Forest added no value: p@10=0.62, p@20=0.61, p@50=0.61, essentially matching the rule. The extra nonlinear capacity earns no gain on near-monotone, low-signal features. This is a valid negative result. The relationship between features and label is approximately linear, and a more complex model does not help.

5. Limitations

This work cannot claim the following:

6. Ranked Recommendations

The action playbook assigns each page a suggested action and reason codes. The top of the queue, ranked by model score:

  1. Review CTR (6,037 pages). These pages have meaningful search demand but low recent CTR. They are the strongest candidates for content refresh. Every page in the model's top 10 was at-risk.
  2. Monitor (9,273 pages). Pages performing adequately. Do not spend editorial time here first.
Action mix chart
Figure 3. Action distribution in the ranked queue. 40% of pages are flagged for CTR review.

Reason codes attached to every pick: recent_search_exposure (impressions >= 200), meaningful_demand (search_volume >= 1,000), transactional_priority.

Reason codes chart
Figure 4. Reason codes across the queue. Recent search exposure and meaningful demand are the most common signals.

Retrain triggers

The model should be re-evaluated when:

7. Reproducibility

All code is in the GitHub repository under work/notebooks/capstone.ipynb. The notebook runs top to bottom with no errors. To reproduce:

  1. Clone the repo
  2. Install dependencies: pip install pandas scikit-learn matplotlib jupyter
  3. Run jupyter notebook work/notebooks/w07_action_playbook.ipynb to generate the feature vector work/outputs/action_playbook_queue.csv
  4. Run jupyter notebook work/notebooks/capstone.ipynb

Key parameters: random_state=42, max_iter=1000, 5-fold GroupKFold by client. The data is loaded from work/outputs/action_playbook_queue.csv, a pre-built feature vector derived from the FlyRank warehouse.

8. Acknowledgments

Data provided by FlyRank AI, a public snapshot hosted on HuggingFace at FlyRank/internship-warehouse. The warehouse contains anonymized search performance data across 104 clients and 519,606 content items. This paper uses a 15,310-item slice with a 60-day window ending August 2025.

Built during the ML Engineering Internship, July - August 2026.

Library Versions::

All seeds set to 42 for reproducibility and CV=5.