Thanks for the great work!
Are there any plans to include LLMs (open source & proprietary) in your benchmark? Something like HELM benchmark would be great. For example, just inference results (with a basic prompt) would be valuable without having to do fine-tuning or in-context learning.
Thanks for the great work!
Are there any plans to include LLMs (open source & proprietary) in your benchmark? Something like HELM benchmark would be great. For example, just inference results (with a basic prompt) would be valuable without having to do fine-tuning or in-context learning.