This comparison includes several popular benchmarks for general language task benchmarks, multimodal task benchmarks and specific downstream task benchmarks. There are 26 benchmarks are compared by total 18 properties to answer different reserach questions such as:
RQ1: what is the data sample size of benchmark?
RQ2: which languages are covered by the benchmark?
RQ3: what are the domain specific tasks of benchmarks?
RQ4: how many models are evaluated on a specific benchmark?
RQ5: how many ability tasks are conducted by benchmark?
ORKG Comparisons have changed. We have added new features and improved the user interface. Comparisons might look slightly different, but the comparison data itself remains unchanged.