Pheebs 是开源 AI 遥测工具,通过 hooks 接入 Claude Code、Cursor、Codex,帮工程团队量化 AI 编程代理的使用情况与协作模式。
热门评论
PH 用户
I've spoken to several engineering leaders who want to better understand their team's AI usage, but don't know where to start. I've also spoken to several engineers who are stuck and don't know how to level up. There's a lot of FOMO building up and being vulnerable about your AI skills is not easy.
Adoption metrics like DAU/WAU and token counts tell you who has access and who uses AI tools. While an ok proxy for AI proficiency and productivity, they say very little about what happens inside a session and how engineers and teams actually use AI.
Which models get used and were they the right fit for the work? Are we spending consciously? Do tests run before a commit? Do outputs get questioned or accepted as is? Do outputs need a lot of repair or do they just get refined? Are repos missing out on scaffolding opportunities (skills, context files)? How's my team doing in AI competencies like context management, evals, and orchestration?
Pheebs captures that and more. It records the shape of sessions and gives engineers a private view of their own work, separate from the team reports. To make sense of the data, Pheebs proposes an AI Proficiency Model, which looks at two things:
1- Repertoire: which harness capabilities show up in an engineer's work across six competencies: models, artifacts, MCP, evals, context management, and orchestration.
2- Judgement (inspired by the Discernment competency from Anthropic's AI Fluency Index): what happens to AI output before a PR is created and merged. Does it get verified, challenged, refined? And was the model used appropriate for the task? Most teams run the largest model for everything. Pheebs shows how much they could save by running smaller models where they fit.
Pheebs is open source so give it a try and send us some feedback!
PH 用户
Hi Pheebs team! I really liked the idea of seeing how often outputs get modified and whether those lead to cost-savings. And also running tests before commits. A few comments / suggestions. 1) Are those metrics shown in the images you have for the Overview? And also, for a given company, is there a way to generate a baseline for a developer's activity that the other developers could follow? for example, maybe you don't need the AI to run and analyze tests after every commit, and only do it before pushing the code? (just thinking out loud for a counter example). Or maybe for a given project, it is better NOT to push back on the code generated by the agent?
2) What do you think of a summary page that shows the recommended steps? You might already have it. I'm only going by the images in the overview.
PH 用户
Moving between subs daily makes it challenging to accurately track use of the models. This seems like it can give a reality check to help encourage more efficient use and model selection.
PH 用户
I enjoy the overall idea. I think a lot of organizations are struggling to figure out how to track this information and ensure that they are getting the value that they thought they would for AI. Does it simply create a report or does it help in creating restrictions, identifying patterns, etc.?
Adoption metrics like DAU/WAU and token counts tell you who has access and who uses AI tools. While an ok proxy for AI proficiency and productivity, they say very little about what happens inside a session and how engineers and teams actually use AI.
Which models get used and were they the right fit for the work? Are we spending consciously? Do tests run before a commit? Do outputs get questioned or accepted as is? Do outputs need a lot of repair or do they just get refined? Are repos missing out on scaffolding opportunities (skills, context files)? How's my team doing in AI competencies like context management, evals, and orchestration?
Pheebs captures that and more. It records the shape of sessions and gives engineers a private view of their own work, separate from the team reports. To make sense of the data, Pheebs proposes an AI Proficiency Model, which looks at two things:
1- Repertoire: which harness capabilities show up in an engineer's work across six competencies: models, artifacts, MCP, evals, context management, and orchestration.
2- Judgement (inspired by the Discernment competency from Anthropic's AI Fluency Index): what happens to AI output before a PR is created and merged. Does it get verified, challenged, refined? And was the model used appropriate for the task? Most teams run the largest model for everything. Pheebs shows how much they could save by running smaller models where they fit.
Pheebs is open source so give it a try and send us some feedback!
2) What do you think of a summary page that shows the recommended steps? You might already have it. I'm only going by the images in the overview.