Population – Earlybirds Invest https://earlybirdsinvest.com Latest Crypto News Thu, 17 Apr 2025 15:12:03 +0000 en-US hourly 1 https://wordpress.org/?v=6.9.7 https://i0.wp.com/earlybirdsinvest.com/wp-content/uploads/2024/12/cropped-New-Project-2024-12-17T235703.455.png?fit=32%2C32&ssl=1 Population – Earlybirds Invest https://earlybirdsinvest.com 32 32 240146708 OpenAI’s o3 scores 136 on Mensa Norway test, surpassing 98% of human population. https://earlybirdsinvest.com/openais-o3-scores-136-on-mensa-norway-test-surpassing-98-of-human-population/ https://earlybirdsinvest.com/openais-o3-scores-136-on-mensa-norway-test-surpassing-98-of-human-population/#respond Thu, 17 Apr 2025 15:12:03 +0000 https://earlybirdsinvest.com/openais-o3-scores-136-on-mensa-norway-test-surpassing-98-of-human-population/

OpenAI’s new “o3” language model achieved an IQ score of 136 on a public Mensa Norway intelligence test, exceeding the threshold for entry into the country’s Mensa chapter for the first time.

The score, calculated from a seven-run rolling average, places the model above approximately 98 percent of the human population, according to a standardized bell-curve IQ distribution used in the benchmarking.

o3 Mensa scores (Source: TrackingAI.org)
o3 Mensa scores (Source: TrackingAI.org)

The finding, disclosed through data from independent platform TrackingAI.org, reinforces the pattern of closed-source, proprietary models outperforming open-source counterparts in controlled cognitive evaluations.

O-series Dominance and Benchmarking Methodology

The “o3” model was released this week and is a part of the “o-series” of large language models, accounting for most top-tier rankings across both test types evaluated by TrackingAI.

The two benchmark formats included a proprietary “Offline Test” curated by TrackingAI.org and a publicly available Mensa Norway test, both scored against a human mean of 100.

While “o3” posted a 116 on the Offline evaluation, it saw a 20-point boost on the Mensa test, suggesting either enhanced compatibility with the latter’s structure or data-related confounds such as prompt familiarity.

The Offline Test included 100 pattern-recognition questions designed to avoid anything that might have appeared in the data used to train AI models.

Both assessments report each model’s result as an average across the seven most recent completions, but no standard deviation or confidence intervals were released alongside the final scores.

The absence of methodological transparency, particularly around prompting strategies and scoring scale conversion, limits reproducibility and interpretability.

Methodology of testing

TrackingAI.org states that it compiles its data by administering a standardized prompt format designed to ensure broad AI compliance while minimizing interpretive ambiguity.

Each language model is presented with a statement followed by four Likert-style response options, Strongly Disagree, Disagree, Agree, Strongly Agree, and is instructed to select one while justifying its choice in two to five sentences.

Responses must be clearly formatted, typically enclosed in bold or asterisks. If a model refuses to answer, the prompt is repeated up to ten times.

The most recent successful response is then recorded for scoring purposes, with refusal events noted separately.

This methodology, refined through repeated calibration across models, aims to provide consistency in comparative assessments while documenting non-responsiveness as a data point in itself.

Performance spread across model types

The Mensa Norway test sharpened the delineation between the truly frontier models, with the o3’s 136 IQ marking a clear lead over the next highest entry.

In contrast, other popular models like GPT-4o scored considerably lower, landing at 95 on Mensa and 64 on Offline, emphasizing the performance gap between this week’s “o3” release and other top models.

Among open-source submissions, Meta’s Llama 4 Maverick was the highest-ranked, posting a 106 IQ on Mensa and 97 on the Offline benchmark.

Most Apache-licensed entries fell within the 60–90 range, reinforcing the current limitations of community-built architectures relative to corporate-backed research pipelines.

Multimodal models see reduced scores and limitations of testing

Notably, models specifically designed to incorporate image input capabilities consistently underperformed their text-only versions. For instance, OpenAI’s “o1 Pro” scored 107 on the Offline test in its text configuration but dropped to 97 in its vision-enabled version.

The discrepancy was more pronounced on the Mensa test, where the text-only variant achieved 122 compared to 86 for the visual version. This suggests that some methods of multimodal pretraining may introduce reasoning inefficiencies that remain unresolved at present.

However, “o3” can also analyze and interpret images to a very high standard, much better than its predecessors, breaking this trend.

Ultimately, IQ benchmarks provide a narrow window into a model’s reasoning capability, with short-context pattern matching offering only limited insights into broader cognitive behavior such as multi-turn reasoning, planning, or factual accuracy.

Additionally, machine test-taking conditions, such as instant access to full prompts and unlimited processing speed, further blur comparisons to human cognition.

The degree to which high IQ scores on structured tests translate to real-world language model performance remains uncertain.

As TrackingAI.org’s researchers acknowledge, even their attempts to avoid training-set leakage do not entirely preclude the possibility of indirect exposure or format generalization, particularly given the lack of transparency around training datasets and fine-tuning procedures for proprietary models.

Independent Evaluators Fill Transparency Gap

Organizations such as LM-Eval, GPTZero, and MLCommons are increasingly relied upon to provide third-party assessments as model developers continue to limit disclosures about internal architectures and training methods.

These “shadow evaluations” are shaping the emerging norms of large language model testing, especially in light of the opaque and often fragmented disclosures from leading AI firms.

OpenAI’s o-series holds a commanding position in this testing workflow, though the long-term implications for general intelligence, agentic behavior, or ethical deployment remain to be addressed in more domain-relevant trials. The IQ scores, while provocative, serve more as signals of short-context proficiency than a definitive indicator of broader capabilities.

Per TrackingAI.org, additional analysis on format-based performance spreads and evaluation reliability will be necessary to clarify the validity of current benchmarks.

With model releases accelerating and independent testing growing in sophistication, comparative metrics may continue to evolve in both format and interpretation.

Mentioned in this article
Posted In: AI, Technology
]]>
https://earlybirdsinvest.com/openais-o3-scores-136-on-mensa-norway-test-surpassing-98-of-human-population/feed/ 0 31309
Dogecoin Shark & Whale Population Rises—Price Turnaround Incoming? https://earlybirdsinvest.com/dogecoin-shark-whale-population-rises-price-turnaround-incoming/ https://earlybirdsinvest.com/dogecoin-shark-whale-population-rises-price-turnaround-incoming/#respond Wed, 19 Mar 2025 11:01:40 +0000 https://earlybirdsinvest.com/dogecoin-shark-whale-population-rises-price-turnaround-incoming/ On-chain data shows the Dogecoin shark and whale wallets have been increasing in number recently, a sign that could be bullish for DOGE’s price.

Dogecoin Sharks & Whales Have Been Expanding Despite Price Decline

According to data from the on-chain analytics firm Santiment, Dogecoin has recently seen a rise in a couple of important indicators. The first metric of relevance here is the “Supply Distribution” of the DOGE wallets carrying more than 1 million tokens.

The Supply Distribution tells us, among other things, the number of addresses that belong to a particular coin range. The indicator for the 1 to 10 coins group, for instance, measures the amount of holders who own at least 1 and at most 10 DOGE in their balance.

The 1 million+ DOGE cohort, which is the range of focus here, includes two key investor groups: sharks and whales. At the current exchange rate, the cutoff for the range converts to around $166,600. This is clearly quite a significant amount, which is why the entities belonging to the sharks and whales are considered important on the network.

Now, here is the chart that shows the trend in the Dogecoin Supply Distribution for the 1 million+ coins range over the last few months:

Dogecoin Supply Distribution

As displayed in the above graph, the Dogecoin Supply Distribution of the sharks and whales observed a plunge when the bearish action in the memecoin’s price first started in January.

Since the start of February, however, the indicator has reversed its direction and has been following an upward trajectory. Interestingly, this wallet increase has come despite the fact that the asset’s decline has only furthered during the period.

The trend would imply that, although the big-money investors panic sold when the drawdown first began, they have since shifted their attention to accumulating the dip instead.

In total, the shark and whale wallets have gone up by 62 (around 1.24%) since the beginning of February and are now not far from the peak witnessed back in January.

The increase in the large wallets isn’t the only positive sign Dogecoin has seen; there has also been bullish development in another indicator attached in the chart. The metric in question is the Active Addresses, which keeps track of the total number of DOGE addresses taking part in some kind of transaction activity on the blockchain every day.

From the graph, it’s visible that the Dogecoin Active Addresses has jumped to a 4-month high recently, suggesting a large amount of users have been making transfers on the network.

While the increase in the shark and whale wallets has been occurring for a while now, the signal in the Active Addresses is a more recent one. It would appear that the current low prices may have finally caught the attention of the masses, who are now coming active to make their moves.

DOGE Price

At the time of writing, Dogecoin is trading around $0.166, up around 4% in the last seven days.

Dogecoin Price Chart

]]>
https://earlybirdsinvest.com/dogecoin-shark-whale-population-rises-price-turnaround-incoming/feed/ 0 25998