The problem with xG totals
There has been several reports in the media suggesting that despite Town’s impressive first half performance against West Ham they were in fact lucky to get a point in the end due to a number of high quality chances given away.
Now when you look at the xG totals from the game it would seem to back this theory up with the xG score being: 1.19-2.04.
The only problem here with these totals is that they take all context out of the individual chances when summing them up, it assumes all chances are independent for the relation to hold. Take Hernandez’s double chance just after half time where he takes a shot from close range which is saved before putting a header over the bar. Now he generates an xG of 0.49 and 0.57 for these two chances respectively, which would suggest that he has a 22% chance of not scoring, 50% chance of scoring once and a 28% chance of scoring twice. Although this can’t be true because how could he have scored two goals in that time there was only 1 second between the two shots and if he scores the first there is clearly no second shot. The two chances are clearly not independent and thus E(X+Y) <> E(X)+E(Y) so a slightly different calculation is needed.
The probability of scoring both is reassigned to scoring just the first because if the first goes in there isn’t a second so it has to happen that only the first is scored. The new probabilities result in:
P(1=m,2=m) = 0.22
P(1=m,2=s) = 0.29
P(1=s,2=m) = 0.49
P(1=s,2=s) = 0
Which give an xG of 0.78 instead of 1.07 which is quite a considerable reduction and certainly gives a different outlook on the two chances. If you do this for the whole game taking any chances taking within 30 seconds of each other (time added to extra time per goal so in theory shortest possible time between goals) and rearranging them as a joint probability distribution you get an xG score of: 1.19-1.74.
And while this still suggests West Ham were the better team and created more the game looks a lot closer and the percentage chance of West Ham winning (taken by modelling each shot or series of shots and chance of each being scored and missed to get probabilities of each score like summer to get result) fell from 58% in the old model to 52% in the new model with context. This means that basically half the time Huddersfield would have gotten at least a point out of this game and thus maybe the result is more deserved than first thought.
In summary by just summing xG important context can be lost from the data and make a team look better than they were it is better to model chances taken in close succession as together and related than independently. Realistically in the current system a team creating lots of chances in quick succession end up equally rated as one that can create the same exact chances over a longer period of time and this is overplaying the first teams ability as if they score the first the later ones definitely don’t happen.
Now when you look at the xG totals from the game it would seem to back this theory up with the xG score being: 1.19-2.04.
The only problem here with these totals is that they take all context out of the individual chances when summing them up, it assumes all chances are independent for the relation to hold. Take Hernandez’s double chance just after half time where he takes a shot from close range which is saved before putting a header over the bar. Now he generates an xG of 0.49 and 0.57 for these two chances respectively, which would suggest that he has a 22% chance of not scoring, 50% chance of scoring once and a 28% chance of scoring twice. Although this can’t be true because how could he have scored two goals in that time there was only 1 second between the two shots and if he scores the first there is clearly no second shot. The two chances are clearly not independent and thus E(X+Y) <> E(X)+E(Y) so a slightly different calculation is needed.
The probability of scoring both is reassigned to scoring just the first because if the first goes in there isn’t a second so it has to happen that only the first is scored. The new probabilities result in:
P(1=m,2=m) = 0.22
P(1=m,2=s) = 0.29
P(1=s,2=m) = 0.49
P(1=s,2=s) = 0
Which give an xG of 0.78 instead of 1.07 which is quite a considerable reduction and certainly gives a different outlook on the two chances. If you do this for the whole game taking any chances taking within 30 seconds of each other (time added to extra time per goal so in theory shortest possible time between goals) and rearranging them as a joint probability distribution you get an xG score of: 1.19-1.74.
And while this still suggests West Ham were the better team and created more the game looks a lot closer and the percentage chance of West Ham winning (taken by modelling each shot or series of shots and chance of each being scored and missed to get probabilities of each score like summer to get result) fell from 58% in the old model to 52% in the new model with context. This means that basically half the time Huddersfield would have gotten at least a point out of this game and thus maybe the result is more deserved than first thought.
In summary by just summing xG important context can be lost from the data and make a team look better than they were it is better to model chances taken in close succession as together and related than independently. Realistically in the current system a team creating lots of chances in quick succession end up equally rated as one that can create the same exact chances over a longer period of time and this is overplaying the first teams ability as if they score the first the later ones definitely don’t happen.
Comments
Post a Comment