在前一篇測試中, 我將工具模式 tools.profile 改成 minimal 觀察 in tokens 數量, 結果從 coding 模式的 21K 驟降為 6.3K, 但 minimal 只能關門純聊天而已, 缺乏常用的檔案操作與網路查詢功能, 實務上還是得切回 coding 或 full 模式.
其次, 我原先的設定檔中模型清單 models 只寫 openai/gpt-5.6 太粗略了, 丟給 OpenAI API 後通常會調用最貴的 gpt-5.6-sol 模型, 其價格為 Input $4/1M, Output $20/1M, 以 21K 低消 token 計算, 每次對話至少 21K × $4/1M≈ $0.084, 約合台幣 2.7 元, 難怪說聲早安就燒掉快 3 元. 而最便宜的 GPT 5.6 模型是 gpt-5.6-luna, 其價格為 Input $0.2/1M, Output $1.2/1M, 同樣的 21K 低消只要 21K × $0.2/1M ≈ $0.0042, 約合台幣 0.14 元, 約是 sol 的 1/20!
本篇旨在測試 openai/gpt-5.6-luna 效能是否符合我的需求.
1. 備份 OpenClaw 設定檔 :
在進行模式變更前先備份設定檔 :
pi@pi4b:~ $ cp ~/.openclaw/openclaw.json ~/.openclaw/openclaw.json.bak
pi@pi4b:~ $ ls -l ~/.openclaw/openclaw.json*
-rw------- 1 pi pi 2917 10月 8 10:08 /home/pi/.openclaw/openclaw.json
-rw------- 1 pi pi 2917 10月 9 00:59 /home/pi/.openclaw/openclaw.json.bak
-rw------- 1 pi pi 2855 10月 8 02:45 /home/pi/.openclaw/openclaw.json.bak.1
-rw------- 1 pi pi 2855 10月 8 02:43 /home/pi/.openclaw/openclaw.json.bak.2
-rw------- 1 pi pi 2917 10月 9 00:57 /home/pi/.openclaw/openclaw.json.bak.202610090057
-rw------- 1 pi pi 2864 10月 8 02:33 /home/pi/.openclaw/openclaw.json.bak.3
-rw------- 1 pi pi 2864 10月 6 23:27 /home/pi/.openclaw/openclaw.json.bak.4
-rw------- 1 pi pi 2917 10月 8 10:20 /home/pi/.openclaw/openclaw.json.last-good
-rw------- 1 pi pi 1455 9月 5 15:58 /home/pi/.openclaw/openclaw.json.manual.bak
2. 修改 tool.profile 為 coding 模式 :
查詢目前工具模式設定為 minimal :
pi@pi4b:~ $ openclaw config get tools.profile
OpenClaw 2026.7.1-2 (0790d9f)
I'll butter your workflow like a lobster roll: messy, delicious, effective.
minimal
更改設定為 coding 模式 :
pi@pi4b:~ $ openclaw config set tools.profile "coding"
│
◇
OpenClaw 2026.7.1-2 (0790d9f)
I'm the reason your shell history looks like a hacker-movie montage.
Updated tools.profile. No gateway restart needed.
確認已改為 coding 模式 :
pi@pi4b:~ $ openclaw config get tools.profile
OpenClaw 2026.7.1-2 (0790d9f)
I've read more man pages than any human should—so you don't have to.
coding
3. 修改 models 的 GPT 5.6 為 luna 模型 :
查詢目前設定 :
pi@pi4b:~ $ openclaw config get agents.defaults.models
OpenClaw 2026.7.1-2 (0790d9f) — I autocomplete your thoughts—just slower and with more API calls.
{
"openai/gpt-5.6": {
"alias": "GPT"
},
"google/gemini-2.5-flash": {
"alias": "Flash"
}
}
此處的 openai/gpt-5.6 太粗略, 用下列指令改為精確的 openai/gpt-5.6-luna :
pi@pi4b:~ $ openclaw config set --replace agents.defaults.models '{"openai/gpt-5.6-luna": {"alias": "GPT"}, "google/gemini-2.5-flash": {"alias": "Flash"}}'
│
◇
OpenClaw 2026.7.1-2 (0790d9f) — Making 'I'll automate that later' happen now.
Updated agents.defaults.models. Change will apply without restarting the gateway.
確認模型已改為 gpt-5.6-luna :
pi@pi4b:~ $ openclaw config get agents.defaults.models
OpenClaw 2026.7.1-2 (0790d9f) — Half butler, half debugger, full crustacean.
{
"google/gemini-2.5-flash": {
"alias": "Flash"
},
"openai/gpt-5.6-luna": {
"alias": "GPT"
}
}
4. 修改 model 的 primary 與 fallback 設定 :
此番主要是想測試 gpt-5.6-luna 模型之效能, 所以要暫時將其設為 primary :
pi@pi4b:~ $ openclaw config set agents.defaults.model.primary "openai/gpt-5.6-luna"
│
◇
OpenClaw 2026.7.1-2 (0790d9f) — Your config is valid, your assumptions are not.
Updated agents.defaults.model.primary. Change will apply without restarting the gateway.
把 gemini-2.5-flash 暫時設為 fallback :
pi@pi4b:~ $ openclaw config set agents.defaults.model.fallbacks '["google/gemini-2.5-flash"]'
│
◇
OpenClaw 2026.7.1-2 (0790d9f)
The only open-source project where the mascot could eat the competition.
Updated agents.defaults.model.fallbacks. Change will apply without restarting the gateway.
檢視設定是否正確 :
pi@pi4b:~ $ openclaw config get agents.defaults.model
OpenClaw 2026.7.1-2 (0790d9f)
Your second brain, except this one actually remembers where you left things.
{
"primary": "openai/gpt-5.6-luna",
"fallbacks": [
"google/gemini-2.5-flash"
]
}
5. 驗證設定檔與重啟服務 :
檢查設定檔 openclaw.json 是否合規 :
pi@pi4b:~ $ openclaw config validate
│
◇
OpenClaw 2026.7.1-2 (0790d9f) — Powered by open source, sustained by spite and good documentation.
Config valid: ~/.openclaw/openclaw.json
用 kill 指令殺死行程來重啟 OpenClaw 服務, 這樣上面的設定才會生效 :
pi@pi4b:~ $ pkill -9 -f openclaw
pi@pi4b:~ $
6. 測試 in tokens 與計算費用 :
在使用 LINE 測試之前, 先查詢 OpenAI API 餘額 :
用 LINE 頻道開啟新對話後問聲早安 :
再次查詢 OpenAI API 餘額 :
用掉 11K tokens, 約 21K 的一半, 吃掉 0.01 美金, 約合台幣 0.3 元, 以上面的價格估計應該是一半 0.07 元台幣才對, 是價格上調了嗎? 但也比 sol 的 2~3 元台幣便宜許多.
檢視狀態 :
確實是使用 gpt-5.6-luna 模型. 但這回卻是 11K in tokens, 可見 GPT 的 tokenizer 計算 token 的數量似乎是 Gemini 的一半左右.
6. 能力測試 :
以下對 Luna 進行五項能力考驗 :
(1). 網路搜尋 :
提示詞如下 :
台積電 ADR 目前價格?
可見已有 34K 快取, 新增 6.4K 快取就是這次搜尋網頁的 token 數量.
(2). Python 知識 :
提示詞如下 :
請用繁體中文解釋 Python 的 list、tuple、set 三者差異。
要求:
1. 用表格比較
2. 各給一個實際例子
3. 說明什麼情況下應該選哪一個
結果如下 :
檢驗表格的內容正確, 三種資料型態的主要差異都有提到, 範例也很生活化, 只有 set 寫得有點不精確, set 是否支援索引它說不可以, 應該說 set 無索引概念較好; 此外 Luna 的回答中也沒有提到時間複雜度. 不過整體來說, Luna 回答結構完整, 沒有明顯錯誤, 範例也合理. 對於 80~90% 的 Python 學習與 Vibe Coding 場景已很夠用, 不需要用到 Terra 或 Sol.
(3). 推理能力測試 :
提示詞如下 :
有 100 顆球,其中 1 顆重量不同,但不知道它比較重還是比較輕。只有一個沒有砝碼的天平,最多可以秤 3 次。請找出異常球,並判斷它是比較重還是比較輕。請完整說明策略。
它的回覆如下 :
這題在 最多 3 次秤重下無解,因此不存在能保證找出異常球的完整策略。
為什麼不可能?
每顆球都有兩種可能:
- 第 1 顆較重
- 第 1 顆較輕
- 第 2 顆較重
- 第 2 顆較輕
- ……
所以共有:
\[
100 \times 2 = 200
\]
種可能情況。
每次使用天平,結果只有三種:
1. 左邊較重
2. 右邊較重
3. 兩邊平衡
因此 3 次秤重最多產生:
\[
3^3 = 27
\]
種不同結果序列。
但我們需要區分 200 種可能,而 3 次秤重最多只能區分 27 種:
\[
27 < 200
\]
所以無論如何安排秤重,都不可能同時:
- 找出是哪一顆球
- 判斷它是較重還是較輕
甚至只找出異常球也不夠
即使不要求判斷重量,只要求找出哪一顆,100 顆球也需要區分 100 種情況;3 次秤重最多只有 27 種結果,仍然不夠。
至少需要幾次?
單純從資訊量來看:
\[
3^4 = 81 < 200
\]
\[
3^5 = 243 \geq 200
\]
因此至少需要 5 次秤重,3 次一定不可能。
結論:這是一道以資訊量為陷阱的題目;正確答案不是設計秤重流程,而是證明 3 次秤重無法解決 100 顆球的問題。
Usage: 493 in / 949 out · cache 11k cached / 916 new
可見 Luna 推理完整, 結論與數學證明正確, 且沒有硬掰策略 (用了資訊理論的解法), 表現非常棒. Luna 在推理能力方面不僅能用而且成本很划算.
(4). Python 程式能力 :
提示詞 :
請寫一個 Python 函式:
def detect_drawdown(prices, threshold=0.2):
輸入是一個每日收盤價 list,找出所有從歷史高點回撤超過 threshold 的區段。
要求:
- 使用 pandas
- 回傳每個 drawdown 的 peak date、trough date、peak price、trough price、drawdown
- 處理空資料
- 提供可直接執行的完整程式
- 附一組測試資料
回應如下 :
以下程式將 drawdown 表示為負百分比,例如 -0.30 代表回撤 30%。 「區段」定義為:從歷史高點開始,直到價格重新回到該高點;若尚未恢復,則截至資料最後一天。
輸出結果類似:
若直接傳入普通 list,peak_date 和 trough_date 會是資料位置,例如 0、3、5。
Usage: ? in / 2.0k out · cache 12k cached / 1.5k new
這表示 Luna 找到了兩個 Drawdown Event, 例如從 2024-01-02 的 110 元高價滑落至 2024-01-05 的70 元, 回撤率為 -36.36%, 結果與數學計算一致, 這顯示程式不僅正確, Drawdown 的算法邏輯也是正確的, 功能完全符合需求. 從這裡可看出 Luna 不只是會寫程式碼, 它還會驗證自己寫的程式碼.
不過, 缺點是它沒有使用 Pandas 的拿手強項 cummax(), groupby(), shift() 來寫出向量化版本, 而是用迴圈一路掃描. 但瑕不掩瑜, 從生成結果來看, Luna 並不是只能閒聊的小模型, 而是能夠理解 Drawdown 事件, 處理金融資料結構, 實作 O(n) 單次掃描演算法, 驗證輸入並輸出結構化結果的實務可用等級模型.
(5). Tool Calling 能力試驗 :
提示詞如下 :
請在目前 Raspberry Pi 上執行:
uname -a
free -h
df -h
然後整理成一個簡短的系統狀態報告。
回應如下 :
結果非常漂亮, 此題並非僅僅是考驗 Luna 的 Tool Calling 能力而已, 而是同時檢驗它的工具結果理解能力與系統管理能力, 回應不僅具體地展示 Pi 4 的 OS 與記憶體大小, 也呈現 SWAP 與分割區資訊, 這表示 Shell Tool 有被成功調用, 指令有真的執行, 且 Luna 有成功地讀回結果並整理成摘要, 表現出色. 另外 Token 成本 3 in / 263 out 也非常低, 從以上能力表現來看, 性價比非常高.















沒有留言 :
張貼留言