This is not because the models are better.
These services have unknown and opaque levels of shadow prompting[1] to tweak the behavior.
The subject article even mentions "tweaking their outputs to the liking of whoever pays the most".
The more I play with LLMs locally,
the more I realize how much prompting going on under the covers is shaping the results from the big tech services.
1 https://www.techpolicy.press/shining-a-light-on-shadow-promp...