Download the dataset from https://allenai.org/data/cord-19
- Download the dataset to the root
./Dataset/which is used in the attribute ofDataProcessorin./data_processor/data_processor.pyandSbertFinetunein./train_model/sbert_finetune.py
Processing dataset to the json and txt files that are easy to process in the following steps.
cd data_processor
python data_processor.py- Training the TFIDF, LDA and fine-tuned sentence-BERT models.
python main.py --phase lda --train_model
python main.py --phase tfidf --train_model
python main.py --phase sbert --train_model --finetune --epoch 1 --loss ContrastiveLoss --sbert_model_name all-mpnet-base-v2- Calculating the indicators of reference support and diversity.
python main.py --phase lda --train_model --similarity
python main.py --phase tfidf --train_model --diversity
python main.py --phase sbert --finetune --epoch 1 --loss ContrastiveLoss --sbert_model_name all-mpnet-base-v2 --similarity
python main.py --phase sbert --finetune --epoch 1 --loss ContrastiveLoss --sbert_model_name all-mpnet-base-v2 --diversitypython ./data_processor/word_frequency.py --count_paper_num
python ./data_processor/word_frequency.py --get_tf # get the word tf in the year and month dimension
python ./data_processor/word_frequency.py --analysis # select the words by 3 different strategies
python ./data_processor/word_frequency.py --viz # visualize the result of word frequencycd draw_pic
python draw_pic.py