顯示具有 Hugging Face 標籤的文章。 顯示所有文章
顯示具有 Hugging Face 標籤的文章。 顯示所有文章

2025年6月27日 星期五

如何在 Hugging Face Spaces 平台佈署 Streamlit 網頁應用程式

在前一篇測試中我們利用 Streamlit 在本機建構了一個 web app 讓使用者輸入 OpenAI 金鑰與提示詞呼叫 Image API 來生圖, 參考 :


本篇旨在將此 web app 佈署至 Hugging Face Spaces 平台上, 作法參考 :


首先登入 Hugging Face Spaces 平台 : 


按右上角的 "+New Space" 鈕 : 




輸入空間名稱, 描述, 授權方式等欄位, 點選 Gradio SDK (之前有 Streamlit SDK 可選, 但最近被拿掉, 有可能是改版或維護而暫時下架, 故此處先用 Gradio SDK, 後續再修改 README.md 等設定), 按底下的 "Create Space" 鈕建立空間 : 





按右上角的 "File" 顯示檔案列表 :




點選 README.md 檔案 :




因為建立空間時點選 Gradio SDK, 故裡面的設定為 Gradio 的版本, 按 "Edit" 鈕進入編輯頁面 : 




將其中 SDK 改為 Streamlit 的版本 :

sdk: streamlit  
sdk_version: 1.35.0  
app_file: app.py  

改好後按左下角的 "Commit changes to main" 鈕提交更新 : 





更新結果頁面出現一個提示, 說 Streamlit 有更新的 v1.46.1 可用, 所以再次按 'edit' 修改版本 : 




改好後同樣按左下角的 "Commit changes to main" 鈕提交更新. 

按 main 後面的空間名稱回到檔案列表頁 :




按右上角 "+Contribute" 鈕點選彈出選單的 "Create a new file" : 




在上方文字框輸入 app.py (這是預設應用程式名稱), 然後將 web app 程式 streamlit_openai_image_api_test_1.py 貼到下方 Edit 輸入框裡面後, 按左下角的 "Commit changes to main" 鈕提交建立檔案 : 



  
同樣按 main 後面的空間名稱回到檔案列表頁, 可見 app.py 已在列表中. 因為此 web app 有用到 Hugging Face Spaces 平台沒有預先安裝的第三方套件 openai, 必須新增一個 requirements.txt 將 openai 列在裡面讓平台去安裝, 按右邊 "+Contribute" 鈕點選彈出選單的 "Create a new file" : 




在上方文字框輸入 requirements.txt, 在下方 Edit 文字框輸入 openai, 按左下角的 "Commit changes to main" 鈕提交建立檔案 : 




這樣就完成全部設定了 (Streamlit 不用列在 requirements.txt 裡面, 因為 Hugging Face Spaces 本身有預先安裝), 按 Spaces 後面的空間名稱即可執行此 web app :




結果如下 :




等 Hugging Face Spaces 平台重新將 Streamlit SDK 上架就不需要去修改 README.md 檔了. 

2025年6月26日 星期四

如何在 Hugging Face Spaces 平台佈署 Gradio 網頁應用程式

Hugging Face Spaces 是AI 開源平台 Hugging Face 提供的免費雲端 Python Web app 主機服務 (也有提供收費方案), 我曾經在測試 Gradio 時紀錄佈署的方法, 參考 :


不過目前網頁似乎有些小更動, 所以下面以昨天用 Gradio 寫的 OpenAI Image API 測試程式為例, 重新寫一篇說明如何將 web app 佈署到 Hugging Face Spaces 平台上. 

首先須註冊 Hugging Face 帳號 : 


登入後點選右上角的 "Spaces" 超連結 : 




按右上角 "+ New Space" 按鈕 : 




填寫 Space Name 與 Description 欄位, 勾選授權方式, 點選 Gradio SDK 後按左下角的 "Create Space" 鈕建立託管空間 :





完成後會自動進入此空間之頁面, 按右上角的三個小點按鈕, 點選彈出選單中的 File 進入檔案管理頁面 : 




按右方的 "+Contribute" 按鈕, 點選彈出選單中的 "Create a new file" : 




在空間名稱後面的文字框輸入 app.py, 這是 web app space 執行時預設會去尋找的主程式名稱 (可以用其他名稱, 但必須先在 README.md 或 .huggingface.yml 指定 app_file 主程式檔名), 然後將 web app 程式貼在下方程式輸入區, 然後按左下方的 "Commit new file to main" 鈕即可 :





這樣便新增了主程式 app.py, 按上方 Spaces 後面的空間名稱即可執行 web app; 按下方 main 後面的專案名稱則會顯示此空間內的檔案列表 : 




但這個 App 有用到平台未預先安裝的套件 openai, 所以直接執行會出現錯誤, 按右上角的三個小點按鈕點選 Files, 這樣也會顯示檔案列表 : 




然後重複上面新增 app.py 的做法, 按右方的 "+Contribute" 按鈕, 點選彈出選單中的 "Create a new file", 在上面檔案名稱文字框輸入 requirements.txt, 在下方 Edit 文字框輸入 openai,  按左下方的 "Commit new file to main" 鈕即可建立此套件安裝檔 : 





這時按 Spaces 後面的空間名稱就可以順利執行此 web app 了 :





這樣便完成 web app 的佈署了 : 


全部 app 列表參考 :


2024年2月2日 星期五

NLP 學習筆記 : 安裝 Hugging Face NLP 工具集套件

Transformer 是目前自然語言處理最常用也最先進的架構, 它可以執行文本的分類 (classification), 產生 (generation), 摘要 (summerization), 與問答 (Q&A) 等任務. Hugging Face 提供了開源的 NLP 工具集套件, 透過統一的介面讓使用者能方便地開發 NLP 應用, 其主要工具集如下 :
  • transformers 套件 : 實作 Transformer 架構
  • datasets 資料集 : 統一的資料集處理工具
參考書籍 :
教學文件參考 :



一. 安裝 transformers 套件 :

今天下午在 Thonny 中 (自帶 Python 3.10) 安裝了 Hugging Face 的 transformers 套件 : 

D:\python>pip install transformers    
Collecting transformers
  Downloading transformers-4.37.2-py3-none-any.whl.metadata (129 kB)
     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 129.4/129.4 kB 692.6 kB/s eta 0:00:00
Requirement already satisfied: filelock in c:\users\tony1\appdata\roaming\python\python310\site-packages (from transformers) (3.12.3)
Requirement already satisfied: huggingface-hub<1.0,>=0.19.3 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from transformers) (0.20.1)
Requirement already satisfied: numpy>=1.17 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from transformers) (1.24.3)
Requirement already satisfied: packaging>=20.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from transformers) (23.1)
Requirement already satisfied: pyyaml>=5.1 in c:\users\tony1\appdata\local\programs\thonny\lib\site-packages (from transformers) (6.0.1)
Collecting regex!=2019.12.17 (from transformers)
  Downloading regex-2023.12.25-cp310-cp310-win_amd64.whl.metadata (41 kB)
     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 42.0/42.0 kB 2.0 MB/s eta 0:00:00
Requirement already satisfied: requests in c:\users\tony1\appdata\roaming\python\python310\site-packages (from transformers) (2.31.0)
Collecting tokenizers<0.19,>=0.14 (from transformers)
  Downloading tokenizers-0.15.1-cp310-none-win_amd64.whl.metadata (6.8 kB)
Collecting safetensors>=0.4.1 (from transformers)
  Downloading safetensors-0.4.2-cp310-none-win_amd64.whl.metadata (3.9 kB)
Requirement already satisfied: tqdm>=4.27 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from transformers) (4.66.1)
Requirement already satisfied: fsspec>=2023.5.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from huggingface-hub<1.0,>=0.19.3->transformers) (2023.9.1)
Requirement already satisfied: typing-extensions>=3.7.4.3 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from huggingface-hub<1.0,>=0.19.3->transformers) (4.9.0)
Requirement already satisfied: colorama in c:\users\tony1\appdata\local\programs\thonny\lib\site-packages (from tqdm>=4.27->transformers) (0.4.6)
Requirement already satisfied: charset-normalizer<4,>=2 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests->transformers) (3.2.0)
Requirement already satisfied: idna<4,>=2.5 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests->transformers) (3.4)
Requirement already satisfied: urllib3<3,>=1.21.1 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests->transformers) (2.1.0)
Requirement already satisfied: certifi>=2017.4.17 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests->transformers) (2023.7.22)
Downloading transformers-4.37.2-py3-none-any.whl (8.4 MB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 8.4/8.4 MB 2.6 MB/s eta 0:00:00
Downloading regex-2023.12.25-cp310-cp310-win_amd64.whl (269 kB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 269.5/269.5 kB 5.5 MB/s eta 0:00:00
Downloading safetensors-0.4.2-cp310-none-win_amd64.whl (269 kB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 269.5/269.5 kB 3.3 MB/s eta 0:00:00
Downloading tokenizers-0.15.1-cp310-none-win_amd64.whl (2.2 MB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2.2/2.2 MB 3.4 MB/s eta 0:00:00
Installing collected packages: safetensors, regex, tokenizers, transformers
Successfully installed regex-2023.12.25 safetensors-0.4.2 tokenizers-0.15.1 transformers-4.37.2

看起來 transformers 套件並沒有很大, 安裝完先匯入整個模組來檢查版本 :

>>> import transformers  
>>> transformers.__version__   
'4.37.2'   

然後用 dir() 檢視其內容 : 

>>> dir(transformers)   
['ALBERT_PRETRAINED_CONFIG_ARCHIVE_MAP', 'ALBERT_PRETRAINED_MODEL_ARCHIVE_LIST', 'ALIGN_PRETRAINED_CONFIG_ARCHIVE_MAP', 'ALIGN_PRETRAINED_MODEL_ARCHIVE_LIST', 'ALL_PRETRAINED_CONFIG_ARCHIVE_MAP', 'ALTCLIP_PRETRAINED_CONFIG_ARCHIVE_MAP', 'ALTCLIP_PRETRAINED_MODEL_ARCHIVE_LIST', 'ASTConfig', 'ASTFeatureExtractor', 'ASTForAudioClassification', 'ASTModel', 'ASTPreTrainedModel', 'AUDIO_SPECTROGRAM_TRANSFORMER_PRETRAINED_CONFIG_ARCHIVE_MAP', 'AUDIO_SPECTROGRAM_TRANSFORMER_PRETRAINED_MODEL_ARCHIVE_LIST', 'AUTOFORMER_PRETRAINED_CONFIG_ARCHIVE_MAP', 'AUTOFORMER_PRETRAINED_MODEL_ARCHIVE_LIST', 'Adafactor', 'AdamW', 'AdamWeightDecay', 'AdaptiveEmbedding', 'AddedToken', 'Agent', 'AlbertConfig', 'AlbertForMaskedLM', 'AlbertForMultipleChoice', 'AlbertForPreTraining', 'AlbertForQuestionAnswering', 'AlbertForSequenceClassification', 'AlbertForTokenClassification', 'AlbertModel', 'AlbertPreTrainedModel', 'AlbertTokenizer', 'AlbertTokenizerFast', 'AlignConfig', 'AlignModel', 'AlignPreTrainedModel', 'AlignProcessor', 'AlignTextConfig', 'AlignTextModel', 'AlignVisionConfig', 'AlignVisionModel', 'AltCLIPConfig', 'AltCLIPModel', 'AltCLIPPreTrainedModel', 'AltCLIPProcessor', 'AltCLIPTextConfig', 'AltCLIPTextModel', 'AltCLIPVisionConfig', 'AltCLIPVisionModel', 'AlternatingCodebooksLogitsProcessor', 'AudioClassificationPipeline', 'AutoBackbone', 'AutoConfig', 'AutoFeatureExtractor', 'AutoImageProcessor', 'AutoModel', 'AutoModelForAudioClassification', 'AutoModelForAudioFrameClassification', 'AutoModelForAudioXVector', 'AutoModelForCTC', 'AutoModelForCausalLM', 'AutoModelForDepthEstimation', 'AutoModelForDocumentQuestionAnswering', 'AutoModelForImageClassification', 'AutoModelForImageSegmentation', 'AutoModelForImageToImage', 'AutoModelForInstanceSegmentation', 'AutoModelForMaskGeneration', 'AutoModelForMaskedImageModeling', 'AutoModelForMaskedLM', 'AutoModelForMultipleChoice', 'AutoModelForNextSentencePrediction', 'AutoModelForObjectDetection', 'AutoModelForPreTraining', 'AutoModelForQuestionAnswering', 'AutoModelForSemanticSegmentation', 'AutoModelForSeq2SeqLM', 'AutoModelForSequenceClassification', 'AutoModelForSpeechSeq2Seq', 'AutoModelForTableQuestionAnswering', 'AutoModelForTextEncoding', 'AutoModelForTextToSpectrogram', 'AutoModelForTextToWaveform', 'AutoModelForTokenClassification', 'AutoModelForUniversalSegmentation', 'AutoModelForVideoClassification', 'AutoModelForVision2Seq', 'AutoModelForVisualQuestionAnswering', 'AutoModelForZeroShotImageClassification', 'AutoModelForZeroShotObjectDetection', 'AutoModelWithLMHead', 'AutoProcessor', 'AutoTokenizer', 'AutoformerConfig', 'AutoformerForPrediction', 'AutoformerModel', 'AutoformerPreTrainedModel', 'AutomaticSpeechRecognitionPipeline', 'AwqConfig', 'AzureOpenAiAgent', 'BARK_PRETRAINED_MODEL_ARCHIVE_LIST', 'BART_PRETRAINED_MODEL_ARCHIVE_LIST', 'BEIT_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BEIT_PRETRAINED_MODEL_ARCHIVE_LIST', 'BERT_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BERT_PRETRAINED_MODEL_ARCHIVE_LIST', 'BIGBIRD_PEGASUS_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BIGBIRD_PEGASUS_PRETRAINED_MODEL_ARCHIVE_LIST', 'BIG_BIRD_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BIG_BIRD_PRETRAINED_MODEL_ARCHIVE_LIST', 'BIOGPT_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BIOGPT_PRETRAINED_MODEL_ARCHIVE_LIST', 'BIT_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BIT_PRETRAINED_MODEL_ARCHIVE_LIST', 'BLENDERBOT_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BLENDERBOT_PRETRAINED_MODEL_ARCHIVE_LIST', 'BLENDERBOT_SMALL_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BLENDERBOT_SMALL_PRETRAINED_MODEL_ARCHIVE_LIST', 'BLIP_2_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BLIP_2_PRETRAINED_MODEL_ARCHIVE_LIST', 'BLIP_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BLIP_PRETRAINED_MODEL_ARCHIVE_LIST', 'BLOOM_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BLOOM_PRETRAINED_MODEL_ARCHIVE_LIST', 'BRIDGETOWER_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BRIDGETOWER_PRETRAINED_MODEL_ARCHIVE_LIST', 'BROS_PRETRAINED_CONFIG_ARCHIVE_MAP', 'BROS_PRETRAINED_MODEL_ARCHIVE_LIST', 'BarkCausalModel', 'BarkCoarseConfig', 'BarkCoarseModel', 'BarkConfig', 'BarkFineConfig', 'BarkFineModel', 'BarkModel', 'BarkPreTrainedModel', 'BarkProcessor', 'BarkSemanticConfig', 'BarkSemanticModel', 'BartConfig', 'BartForCausalLM', 'BartForConditionalGeneration', 'BartForQuestionAnswering', 'BartForSequenceClassification', 'BartModel', 'BartPreTrainedModel', 'BartPretrainedModel', 'BartTokenizer', 'BartTokenizerFast', 'BarthezTokenizer', 'BarthezTokenizerFast', 'BartphoTokenizer', 'BasicTokenizer', 'BatchEncoding', 'BatchFeature', 'BeamScorer', 'BeamSearchScorer', 'BeitBackbone', 'BeitConfig', 'BeitFeatureExtractor', 'BeitForImageClassification', 'BeitForMaskedImageModeling', 'BeitForSemanticSegmentation', 'BeitImageProcessor', 'BeitModel', 'BeitPreTrainedModel', 'BertConfig', 'BertForMaskedLM', 'BertForMultipleChoice', 'BertForNextSentencePrediction', 'BertForPreTraining', 'BertForQuestionAnswering', 'BertForSequenceClassification', 'BertForTokenClassification', 'BertGenerationConfig', 'BertGenerationDecoder', 'BertGenerationEncoder', 'BertGenerationPreTrainedModel', 'BertGenera…

從後面的 ... 可知其成員非常龐大啊! 

谷歌 Colab 執行環境已經預先安裝 transformers 套件可直接匯入使用, 不需安裝 : 




可見其版本較上面本機安裝的 v4.37.2 要舊一些 (v4.35.2). 


二. 安裝 datasets 工具集套件 :

此工具集負責資料預處理, 例如資料集的載入與儲存, 檢視, 排序, 過濾, 抽樣, 與拆分等, Hugging Face 的 datasets 套件提供了統一的資料集處理 API 介面, 可以大幅降低資料預處理的難度與工作量 (在資料科學中, 預處理佔據大部分工作量, 可說是一種 dirty work). 

此套件在本機可直接用 pip install 安裝 :

D:\python>pip install datasets    
Collecting datasets
  Downloading datasets-2.16.1-py3-none-any.whl.metadata (20 kB)
Requirement already satisfied: filelock in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (3.12.3)
Requirement already satisfied: numpy>=1.17 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (1.24.3)
Requirement already satisfied: pyarrow>=8.0.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (13.0.0)
Collecting pyarrow-hotfix (from datasets)
  Downloading pyarrow_hotfix-0.6-py3-none-any.whl.metadata (3.6 kB)
Requirement already satisfied: dill<0.3.8,>=0.3.0 in c:\users\tony1\appdata\local\programs\thonny\lib\site-packages (from datasets) (0.3.7)
Requirement already satisfied: pandas in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (2.0.3)
Requirement already satisfied: requests>=2.19.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (2.31.0)
Requirement already satisfied: tqdm>=4.62.1 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (4.66.1)
Collecting xxhash (from datasets)
  Downloading xxhash-3.4.1-cp310-cp310-win_amd64.whl.metadata (12 kB)
Collecting multiprocess (from datasets)
  Downloading multiprocess-0.70.16-py310-none-any.whl.metadata (7.2 kB)
Requirement already satisfied: fsspec<=2023.10.0,>=2023.1.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from fsspec[http]<=2023.10.0,>=2023.1.0->datasets) (2023.9.1)
Requirement already satisfied: aiohttp in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (3.9.1)
Requirement already satisfied: huggingface-hub>=0.19.4 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (0.20.1)
Requirement already satisfied: packaging in c:\users\tony1\appdata\roaming\python\python310\site-packages (from datasets) (23.1)
Requirement already satisfied: pyyaml>=5.1 in c:\users\tony1\appdata\local\programs\thonny\lib\site-packages (from datasets) (6.0.1)
Requirement already satisfied: attrs>=17.3.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from aiohttp->datasets) (23.1.0)
Requirement already satisfied: multidict<7.0,>=4.5 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from aiohttp->datasets) (6.0.4)
Requirement already satisfied: yarl<2.0,>=1.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from aiohttp->datasets) (1.9.4)
Requirement already satisfied: frozenlist>=1.1.1 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from aiohttp->datasets) (1.4.1)
Requirement already satisfied: aiosignal>=1.1.2 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from aiohttp->datasets) (1.3.1)
Requirement already satisfied: async-timeout<5.0,>=4.0 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from aiohttp->datasets) (4.0.3)
Requirement already satisfied: typing-extensions>=3.7.4.3 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from huggingface-hub>=0.19.4->datasets) (4.9.0)
Requirement already satisfied: charset-normalizer<4,>=2 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests>=2.19.0->datasets) (3.2.0)
Requirement already satisfied: idna<4,>=2.5 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests>=2.19.0->datasets) (3.4)
Requirement already satisfied: urllib3<3,>=1.21.1 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests>=2.19.0->datasets) (2.1.0)
Requirement already satisfied: certifi>=2017.4.17 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from requests>=2.19.0->datasets) (2023.7.22)
Requirement already satisfied: colorama in c:\users\tony1\appdata\local\programs\thonny\lib\site-packages (from tqdm>=4.62.1->datasets) (0.4.6)
INFO: pip is looking at multiple versions of multiprocess to determine which version is compatible with other requirements. This could take a while.
  Downloading multiprocess-0.70.15-py310-none-any.whl.metadata (7.2 kB)
Requirement already satisfied: python-dateutil>=2.8.2 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from pandas->datasets) (2.8.2)
Requirement already satisfied: pytz>=2020.1 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from pandas->datasets) (2023.3)
Requirement already satisfied: tzdata>=2022.1 in c:\users\tony1\appdata\roaming\python\python310\site-packages (from pandas->datasets) (2023.3)
Requirement already satisfied: six>=1.5 in c:\users\tony1\appdata\local\programs\thonny\lib\site-packages (from python-dateutil>=2.8.2->pandas->datasets) (1.16.0)
Downloading datasets-2.16.1-py3-none-any.whl (507 kB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 507.1/507.1 kB 2.0 MB/s eta 0:00:00
Downloading multiprocess-0.70.15-py310-none-any.whl (134 kB)
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 134.8/134.8 kB 2.0 MB/s eta 0:00:00
Downloading pyarrow_hotfix-0.6-py3-none-any.whl (7.9 kB)
Downloading xxhash-3.4.1-cp310-cp310-win_amd64.whl (29 kB)
Installing collected packages: xxhash, pyarrow-hotfix, multiprocess, datasets
Successfully installed datasets-2.16.1 multiprocess-0.70.15 pyarrow-hotfix-0.6 xxhash-3.4.1

檢查版本 :

>>> import datasets  
>>> datasets.__version__   
'2.16.1'

用 dir() 檢視其內容 :

>>> dir(datasets)     
['Array2D', 'Array3D', 'Array4D', 'Array5D', 'ArrowBasedBuilder', 'Audio', 'AudioClassification', 'AutomaticSpeechRecognition', 'BeamBasedBuilder', 'BuilderConfig', 'ClassLabel', 'Dataset', 'DatasetBuilder', 'DatasetDict', 'DatasetInfo', 'DownloadConfig', 'DownloadManager', 'DownloadMode', 'Features', 'GeneratorBasedBuilder', 'Image', 'ImageClassification', 'IterableDataset', 'IterableDatasetDict', 'LanguageModeling', 'Metric', 'MetricInfo', 'NamedSplit', 'NamedSplitAll', 'QuestionAnsweringExtractive', 'ReadInstruction', 'Sequence', 'Split', 'SplitBase', 'SplitDict', 'SplitGenerator', 'SplitInfo', 'StreamingDownloadManager', 'SubSplitInfo', 'Summarization', 'TaskTemplate', 'TextClassification', 'Translation', 'TranslationVariableLanguages', 'Value', 'VerificationMode', 'Version', '__builtins__', '__cached__', '__doc__', '__file__', '__loader__', '__name__', '__package__', '__path__', '__spec__', '__version__', 'are_progress_bars_disabled', 'arrow_dataset', 'arrow_reader', 'arrow_writer', 'builder', 'combine', 'concatenate_datasets', 'config', 'data_files', 'dataset_dict', 'deprecation_utils', 'disable_caching', 'disable_progress_bar', 'disable_progress_bars', 'doc_utils', 'download', 'enable_caching', 'enable_progress_bar', 'enable_progress_bars', 'exceptions', 'experimental', 'extract', 'features', 'file_utils', 'filesystems', 'fingerprint', 'formatting', 'get_dataset_config_info', 'get_dataset_config_names', 'get_dataset_default_config_name', 'get_dataset_infos', 'get_dataset_split_names', 'hub', 'info', 'info_utils', 'inspect', 'inspect_dataset', 'inspect_metric', 'interleave_datasets', 'is_caching_enabled', 'is_progress_bar_enabled', 'iterable_dataset', 'keyhash', 'list_datasets', 'list_metrics', 'load', 'load_dataset', 'load_dataset_builder', 'load_from_disk', 'load_metric', 'logging', 'metadata', 'metric', 'naming', 'packaged_modules', 'parallel', 'patching', 'percent', 'py_utils', 'search', 'set_caching_enabled', 'sharding', 'splits', 'stratify', 'streaming', 'table', 'tasks', 'tf_utils', 'tqdm', 'track', 'typing', 'utils', 'version']

不過 Colab 並未預載 datasets 套件, 每次重新連線都必須自行安裝一次 :

!pip install datasets 




匯入檢視版本 :




可見與本機一樣都是最新的 v2.16.1 版.