As Large Language Models (LLM) continue to develop, their coding capabilities are increasingly being applied to specific scenarios in industrial production, such as AI tasks that require writing PyTorch code. However, the evaluation of this code's quality necessitates the setup of complex execution environments, which presents numerous challenges for code assessment. Our project, an innovative LLM evaluation system, is specifically designed to assess the code abilities of LLMs in complex, real-world coding environments. In order to tackle the intricate environmental demands, we have strategically utilized Docker containers to simulate the code running environment. Furthermore, to enhance operational efficiency, we have implemented task parallelism using Spark MapReduce and AWS EC2 clusters. By amalgamating the publicly available Humaneval dataset with our own custom-designed machine learning problems, we have integrated these test cases into our evaluation system, thereby creating a robust benchmark. We conducted experiments on several LLMs with our benchmark, the results indicate that GPT-4 outperforms Wenxin Yiyan, Zhipu and GPT-3.5 Turbo, which aligns with user experiences.
This benchmark can help to evaluate LLM's code capability in higher level like CV and NLP tasks. Here is an example case: https://github.com/MichaelYang-lyx/LLM-Code-Benchmark/tree/main/data/AItest/test2.
Our evaluation system encompasses three primary steps to ensure a comprehensive evaluation of each language model's capabilities:
- Infer
- Evaluate
- Summarize
- Docker is enough
To prepare the environment for code testing, you can either build the Docker image by:
docker build -t mypytorch .or pull it from Docker Hub using the commands below:
sudo docker pull michaelyang0050/llm_benchmark
docker tag michaelyang0050/llm_benchmark mypytorchThen you need to create an .env file containning your API keys with the following format:
OPENAI_API_KEY= ...
BAIDU_SECRET_KEY= ...
...
Within the jobs directory, you'll find multiple executable jobs. Each job is tailored to test a different model on a designated dataset, crafted to rigorously assess model performance and precision.
Navigate to the jobs directory and execute the job file of your choice to run a test. Ensure you have the correct permissions and that the environment variables are properly configured.just run:
sudo bash run_cluster.shor if you want to run it locally:
sudo bash run.shWe test four different LLMs based on our benchmark. And the results are as follows:
This project is licensed under the MIT License - see the LICENSE.md file for details.
- The PySpark community
- Docker Hub
- All the contributors who have played a role in developing this project
Thanks goes to these wonderful people (emoji key):
Michael 💻 |
ZixinMa27 💻 |
Z.Shen 💻 |
JifengCHN 💻 |
|||
|
|
||||||
This project follows the all-contributors specification. Contributions of any kind welcome!





