devstats is really useful and we see it leveraged for all sorts of things ... including identifying inactive contributors, celebrating most active contributors, etc.
But as I understand it, the gh archive dataset is missing a LOT of events this year. We see this in Kubernetes when we go to seed the active-contributors list for the purposes of the upcoming steering election, there's a sharp drop off in people above the 50 contributions threshold, which seems to trace back to gh archive missing events.
Please put a banner or something of similar nature to warn potential users that the underlying public dataset is undercounting contributions and should not be treated as 100% accurate.
We know people are making decisions based on this data and we should warn that it undercounts due to the public data source being leveraged missing events.
cc @mrbobbytables
devstats is really useful and we see it leveraged for all sorts of things ... including identifying inactive contributors, celebrating most active contributors, etc.
But as I understand it, the gh archive dataset is missing a LOT of events this year. We see this in Kubernetes when we go to seed the active-contributors list for the purposes of the upcoming steering election, there's a sharp drop off in people above the 50 contributions threshold, which seems to trace back to gh archive missing events.
Please put a banner or something of similar nature to warn potential users that the underlying public dataset is undercounting contributions and should not be treated as 100% accurate.
We know people are making decisions based on this data and we should warn that it undercounts due to the public data source being leveraged missing events.
cc @mrbobbytables