Every Attribution Model Explained With Code
Attribution models with code: first, last, linear, time-decay, Markov, Shapley.

The Problem
Your customer journey has 8 touchpoints across 4 channels over 3 weeks. Your\
CEO wants to know: "Which channel gets credit for this $5,000 deal?"
The answer depends entirely on which attribution model you use. And most people\
using attribution models don't understand what their model is actually doing.
This is every major attribution model, explained with math and code.
The Setup
# Example journey
journey = {
'touchpoints': [
{'channel': 'LinkedIn', 'date': '2025-01-01', 'cost': 50},
{'channel': 'Google', 'date': '2025-01-05', 'cost': 30},
{'channel': 'Email', 'date': '2025-01-10', 'cost': 5},
{'channel': 'Direct', 'date': '2025-01-15', 'cost': 0},
{'channel': 'Google', 'date': '2025-01-18', 'cost': 40},
{'channel': 'Retargeting', 'date': '2025-01-20', 'cost': 25},
{'channel': 'Direct', 'date': '2025-01-21', 'cost': 0},
],
'conversion_value': 5000,
'conversion_date': '2025-01-22'
}
Single-Touch Models
First-Touch Attribution
Logic: 100% credit to the first touchpoint.
def first_touch(journey):
credit = {t['channel']: 0 for t in journey['touchpoints']}
first = journey['touchpoints'][0]['channel']
credit[first] = journey['conversion_value']
return credit
# Result: {'LinkedIn': 5000, 'Google': 0, 'Email': 0, 'Direct': 0, 'Retargeting': 0}
When to use: Measuring top-of-funnel awareness. Understanding which\
channels introduce customers to your brand.
Problem: Ignores everything that happens after first touch. A customer\
who discovered you on LinkedIn but was nurtured through 6 email campaigns\
and 3 Google searches gets credited entirely to LinkedIn.
Last-Touch Attribution
Logic: 100% credit to the final touchpoint before conversion.
def last_touch(journey):
credit = {t['channel']: 0 for t in journey['touchpoints']}
last = journey['touchpoints'][-1]['channel']
credit[last] = journey['conversion_value']
return credit
# Result: {'LinkedIn': 0, 'Google': 0, 'Email': 0, 'Direct': 5000, 'Retargeting': 0}
When to use: Optimizing bottom-funnel conversion. Understanding what\
closes deals.
Problem: Overvalues branded search and direct traffic. Every journey ends\
with "Direct" or "Branded Search" if the user types your URL.
Multi-Touch Models
Linear Attribution
Logic: Equal credit to all touchpoints.
def linear(journey):
n = len(journey['touchpoints'])
share = journey['conversion_value'] / n
credit = {}
for t in journey['touchpoints']:
credit[t['channel']] = credit.get(t['channel'], 0) + share
return credit
# Result: LinkedIn=714, Google=1428, Email=714, Direct=1428, Retargeting=714
When to use: Long sales cycles where every touch matters equally.
Problem: A $50 LinkedIn impression gets the same credit as a $200 Google\
click. Doesn't account for effort or cost.
Time-Decay Attribution
Logic: Credit decays exponentially by time since touchpoint.
import math
def time_decay(journey, half_life_days=7):
conversion_date = datetime.strptime(journey['conversion_date'], '%Y-%m-%d')
weights = []
for t in journey['touchpoints']:
touch_date = datetime.strptime(t['date'], '%Y-%m-%d')
days_before = (conversion_date - touch_date).days
weight = math.pow(0.5, days_before / half_life_days)
weights.append(weight)
total_weight = sum(weights)
credit = {}
for t, w in zip(journey['touchpoints'], weights):
channel = t['channel']
credit[channel] = credit.get(channel, 0) + (w / total_weight) * journey['conversion_value']
return credit
# With 7-day half-life:
# LinkedIn (21 days): weight=0.125
# Google (17 days): weight=0.189
# Email (12 days): weight=0.297
# Direct (7 days): weight=0.5
# Google (4 days): weight=0.67
# Retargeting (2 days): weight=0.82
# Direct (1 day): weight=0.91
#
# Result: LinkedIn=312, Google=937, Email=741, Direct=1406, Retargeting=1152
When to use: Short sales cycles where recency matters.
Problem: Completely ignores the role of awareness. The first touch that\
made the customer aware of you gets almost no credit.
Position-Based (U-Shaped)
Logic: 40% first touch, 40% last touch, 20% split among middle.
def position_based(journey):
n = len(journey['touchpoints'])
credit = {t['channel']: 0 for t in journey['touchpoints']}
first = journey['touchpoints'][0]['channel']
last = journey['touchpoints'][-1]['channel']
credit[first] += journey['conversion_value'] * 0.4
credit[last] += journey['conversion_value'] * 0.4
if n > 2:
middle_share = (journey['conversion_value'] * 0.2) / (n - 2)
for t in journey['touchpoints'][1:-1]:
credit[t['channel']] += middle_share
return credit
# Result: LinkedIn=2000, Google=685, Email=285, Direct=2000, Retargeting=285
When to use: When you care about both acquisition and conversion.
Problem: Arbitrary 40/40/20 split. Why not 30/30/40? The weights are made up.
W-Shaped
Logic: 30% first touch, 30% last touch, 30% opportunity creation, 10% middle.
def w_shaped(journey, opportunity_touch_index=3):
n = len(journey['touchpoints'])
credit = {t['channel']: 0 for t in journey['touchpoints']}
first = journey['touchpoints'][0]['channel']
last = journey['touchpoints'][-1]['channel']
opp = journey['touchpoints'][opportunity_touch_index]['channel']
credit[first] += journey['conversion_value'] * 0.3
credit[last] += journey['conversion_value'] * 0.3
credit[opp] += journey['conversion_value'] * 0.3
remaining = journey['conversion_value'] * 0.1
if n > 3:
others = [t for i, t in enumerate(journey['touchpoints'])
if i not in [0, opportunity_touch_index, n-1]]
share = remaining / len(others)
for t in others:
credit[t['channel']] += share
return credit
When to use: B2B with clear opportunity stages.
Algorithmic Models
Markov Chain Attribution
Logic: Model transition probabilities between channels. Credit channels\
based on their removal effect (conversion rate drops when channel is removed).
import numpy as np
from collections import defaultdict
def build_transition_matrix(journeys):
"""Build transition probability matrix from all journeys."""
transitions = defaultdict(lambda: defaultdict(int))
channel_counts = defaultdict(int)
for journey in journeys:
touchpoints = ['START'] + [t['channel'] for t in journey['touchpoints']] + ['CONVERT']
for i in range(len(touchpoints) - 1):
from_state = touchpoints[i]
to_state = touchpoints[i + 1]
transitions[from_state][to_state] += 1
channel_counts[from_state] += 1
# Build matrix
channels = list(channel_counts.keys()) + ['CONVERT', 'DROP']
n = len(channels)
matrix = np.zeros((n, n))
for i, from_ch in enumerate(channels):
if from_ch in ['CONVERT', 'DROP']:
matrix[i][i] = 1.0 # Absorbing states
continue
total = sum(transitions[from_ch].values())
for j, to_ch in enumerate(channels):
if to_ch in transitions[from_ch]:
matrix[i][j] = transitions[from_ch][to_ch] / total
return matrix, channels
def removal_effect(matrix, channels, target_channel):
"""Calculate conversion rate with channel removed."""
# Set all transitions from target to go to DROP instead
target_idx = channels.index(target_channel)
modified = matrix.copy()
for j in range(len(channels)):
if channels[j] == 'DROP':
modified[target_idx][j] += modified[target_idx][j]
modified[target_idx][j] = 0
# Calculate absorption probability
# (Simplified - full implementation requires solving linear system)
return calculate_conversion_rate(modified)
def markov_attribution(journeys):
matrix, channels = build_transition_matrix(journeys)
base_rate = calculate_conversion_rate(matrix)
effects = {}
for ch in channels:
if ch in ['START', 'CONVERT', 'DROP']:
continue
without_rate = removal_effect(matrix, channels, ch)
effects[ch] = base_rate - without_rate
# Normalize to conversion value
total_effect = sum(effects.values())
attribution = {ch: (eff / total_effect) * total_conversion_value
for ch, eff in effects.items()}
return attribution
When to use: When you have enough journey data (1000+ conversions) and\
want data-driven allocation.
Problem: Computationally intensive. Requires complete journey data.\
Sensitive to data quality.
Shapley Value
Logic: From cooperative game theory. Credit each channel based on its\
marginal contribution across all possible subsets.
from itertools import combinations
def shapley_value(journeys, channels):
"""Calculate Shapley value for each channel."""
# Simplified: estimate from coalition values
coalition_values = estimate_coalition_values(journeys)
n = len(channels)
shapley = {ch: 0 for ch in channels}
for ch in channels:
for coalition in all_coalitions_without(channels, ch):
s = len(coalition)
weight = (math.factorial(s) * math.factorial(n - s - 1)) / math.factorial(n)
v_with = coalition_values.get(tuple(sorted(coalition + [ch])), 0)
v_without = coalition_values.get(tuple(sorted(coalition)), 0)
shapley[ch] += weight * (v_with - v_without)
return shapley
When to use: When channels interact (e.g., Email + Retargeting together\
perform better than either alone).
Problem: Exponentially complex. Impractical for >10 channels.
Incrementality: The Gold Standard
No attribution model tells you causality. Only experiments do.
Geo-Lift Test
def geo_lift_test(treatment_regions, control_regions, pre_period, post_period):
"""
treatment_regions: list of region IDs exposed to campaign
control_regions: list of similar regions not exposed
"""
# Calculate pre-period similarity
pre_treatment = get_sales(treatment_regions, pre_period)
pre_control = get_sales(control_regions, pre_period)
# Scaling factor to match pre-period levels
scale = pre_treatment.mean() / pre_control.mean()
# Post-period
post_treatment = get_sales(treatment_regions, post_period)
post_control = get_sales(control_regions, post_period) * scale
lift = (post_treatment.mean() - post_control.mean()) / post_control.mean()
# Statistical significance
t_stat, p_value = ttest_ind(post_treatment, post_control)
return {
'lift': lift,
'p_value': p_value,
'significant': p_value < 0.05
}
Conversion Lift (Platform Holdout)
def conversion_lift(spend, treatment_conversions, control_conversions,
treatment_size, control_size):
"""
Facebook/Google native lift test results.
"""
treatment_rate = treatment_conversions / treatment_size
control_rate = control_conversions / control_size
incremental_rate = treatment_rate - control_rate
incremental_conversions = incremental_rate * treatment_size
iCPA = spend / incremental_conversions
iROAS = (incremental_conversions * aov - spend) / spend
return {
'incremental_conversions': incremental_conversions,
'incremental_rate': incremental_rate,
'iCPA': iCPA,
'iROAS': iROAS
}
Model Comparison
| Model | Complexity | Data Required | Best For | Key Weakness |
|---|---|---|---|---|
| First-Touch | Low | Any | Awareness | Ignores nurture |
| Last-Touch | Low | Any | Conversion | Overvalues bottom-funnel |
| Linear | Low | Any | Long cycles | No effort weighting |
| Time-Decay | Medium | Any | Short cycles | Ignores first touch |
| Position-Based | Low | Any | B2B | Arbitrary weights |
| Markov Chain | High | 1000+ journeys | Data-driven | Computationally heavy |
| Shapley | Very High | Complete data | Channel interactions | Exponential complexity |
| Incrementality | Medium | Experiment budget | Causality | Expensive, limited scale |
Recommendation
- Start with Position-Based (U-Shaped) for B2B, Time-Decay for B2C
- Validate with incrementality tests quarterly
- Move to Markov Chain when you have 1000+ complete journeys
- Never use a single model for all decisions
- Report confidence intervals, not just point estimates
The right attribution model is the one that helps you make better decisions,\
not the one with the most math.