Sure. The topic isn’t really Stack Exchange. But now that it went off topic and you added the “10 seconds” comment: Did you look at the relative volume of STEM vs non-STEM questions/answers on stackexchange vs the other sites? [Hint: For Stack Exchange have a look at the question volumes by “topic”: https://data.stackexchange.com/ ]
Everybody seems to miss my point. My overall point/hypothesis is that the results of what Semrush collected has more to do with their question database than it has to do with the source of the LLMs training data. It’s definitely a flawed conclusion. And I provided an example to illustrate why that’s the case.
If you ask the right question, and the question is not STEM, you could still easily hit StackExchange.
Sometimes when the search engine throws up StackExchange as the solution to my STEM problem, I will notice some left-field (non-STEM) question over on the right hand side of the window. Occasionally I will even click to read.
Today’s random question: blowing your nose during krias shema and shemoneh esrei (I didn’t click that though because I don’t even understand the question. But no need for anyone to digress onto that to explain it as this is not Round Table.)
Very little of StackExchange’s content is not STEM. So, no, not “easily” and certainly it would not be “cited much”.
And, my point again: Semrush’s conclusion is not valid. That data can’t be used to make a conclusion about how much of the LLM’s training set consisted of which source. I hypothesize that Semrush’s results say much more about where Semrush got its database of queries.