The original study itself had at least one developer who later revealed that he had filtered out tasks he prefered not to do without AI: https://xcancel.com/ruben_bloom/status/1943536052037390531 -- given the N was 16, and he seems to have been one of the more AI-experienced devs, and we don't know if the other devs did this, the results of the first study itself could be questioned.
Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions.
I am not at all suggesting the first study is good, or that I believe its conclusions.
(Or that the failure of the second study validates the conclusions of the first.)
I am just saying that people here who think the second study overturned, debunked or corrected the findings of the first are explicitly wrong, because even its authors admit it is a broken study.
It would take a non-broken study to do that, and it may not actually be possible anymore, which is perhaps the most useful finding of the second study.