Skip to content

feat(route/huggingface): add user posts activity - #23262

Open
qqc20043 wants to merge 3 commits into
DIYgod:masterfrom
qqc20043:feat/route-huggingface-user-posts
Open

qqc20043 wants to merge 3 commits into
DIYgod:masterfrom
qqc20043:feat/route-huggingface-user-posts

Conversation

@qqc20043

@qqc20043 qqc20043 commented Sep 11, 2026 •

Copy link
Copy Markdown

Involved Issue / 该 PR 相关 Issue

Close #23059

Example for the Proposed Route(s) / 路由地址示例

/huggingface/activity/AbstractPhil/posts

New RSS Route Checklist / 新 RSS 路由检查表

  • New Route / 新的路由
  • Documentation / 文档说明
  • Full text / 全文获取
  • Use cache / 使用缓存
  • Anti-bot or rate limit / 反爬/频率限制
    • If yes, do your code reflect this sign? / 如果有, 是否有对应的措施?
  • Date and time / 日期和时间
    • Parsed / 可以解析
    • Correct time zone / 时区正确
  • New package added / 添加了新的包
  • Puppeteer

Note / 说明

新增 Hugging Face 用户帖子动态路由,补齐 huggingface namespace 下缺失的 posts 类型
(现有 blog、blog-community、daily-papers、models、user-likes 均无用户帖子)。

实现说明:

  1. 数据来源:Hugging Face 未提供按用户过滤的公开帖子 API
    (/api/posts?author= 的 author 参数被服务端忽略,始终返回全站最新条目;
    /api/users/{user}/posts 返回 404;/api/social-posts 需要认证)。
    因此本路由读取 https://huggingface.co/{user}/activity/posts 页面,
    该页为 SvelteKit 服务端渲染,帖子数据完整嵌入在
    data-target="UserProfile" 节点的 data-props 属性中,无需浏览器渲染即可提取。

  2. 条目数量:服务端渲染的首屏数据包含最新的 2 条帖子(?p= 等分页参数对该页无效),
    路由直接输出这 2 条。totalPosts 字段仅用于参考,未作为分页依据。

  3. 全文输出:rawContent 为帖子原始 Markdown,经 markdown-it 渲染为 HTML 作为
    description,因此订阅端可直接获得全文与内嵌图片。

  4. 反爬:Hugging Face 部署了 AWS WAF,短时间高频请求会返回 429。
    已将 features.antiCrawler 置为 true,RSSHub 的请求层自带重试与退避。

  5. 时间:publishedAt 为 ISO 8601 UTC 时间,统一交由 parseDate 处理。

本地验证(示例用户 AbstractPhil):

  • /huggingface/activity/AbstractPhil/posts → 2 条,标题、链接、pubDate、作者均正确
  • 链接形如 https://huggingface.co/posts/AbstractPhil/372092347372730

@github-actions github-actions Bot added the route label Sep 11, 2026
Comment thread lib/routes/huggingface/user-posts.ts Fixed
@github-actions

Copy link
Copy Markdown
Contributor

Insufficient data

Not enough activity yet to make a reliable assessment.

View full analysis →

This is an automated analysis by AgentScan

@github-actions

github-actions Bot commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Auto Review

No clear rule violations found in the current diff.

@github-actions

Copy link
Copy Markdown
Contributor

Auto Route Test failed, please check your PR body format and reopen pull request. Check logs for more details.
自动路由测试失败,请确认 PR 正文部分符合格式规范并重新开启,详情请检查 日志。

@github-actions github-actions Bot added the auto: route no found Automated test failed due to route can not be found in PR description body label Sep 11, 2026
@github-actions github-actions Bot closed this Sep 11, 2026
@github-actions github-actions Bot added the auto: DO NOT merge Docker image won't even start label Sep 11, 2026
@github-actions github-actions Bot reopened this Sep 11, 2026
@github-actions github-actions Bot removed the auto: DO NOT merge Docker image won't even start label Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Please use actual values in routes section instead of path parameters.
请在 routes 部分使用实际值而不是路径参数。

@github-actions github-actions Bot removed the auto: route no found Automated test failed due to route can not be found in PR description body label Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Successfully generated as following:

http://localhost:1200/huggingface/activity/AbstractPhil/posts - Success ✔️
<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>AbstractPhil - Posts Activity</title>
    <link>https://huggingface.co/AbstractPhil/activity/posts</link>
    <atom:link href="http://localhost:1200/huggingface/activity/AbstractPhil/posts" rel="self" type="application/rss+xml"></atom:link>
    <description>AbstractPhil - Posts Activity - Powered by RSSHub</description>
    <generator>RSSHub</generator>
    <webMaster>contact@rsshub.app (RSSHub)</webMaster>
    <language>en</language>
    <lastBuildDate>Fri, 11 Sep 2026 08:08:11 GMT</lastBuildDate>
    <ttl>5</ttl>
    <item>
      <title>The post-beatrix-2s and control variant article is finally satisfactory, so the article is now released https://huggingface.co/blog/AbstractPhil/beatr...</title>
      <description>&lt;p&gt;The post-beatrix-2s and control variant article is finally satisfactory, so the article is now released &lt;a href=&quot;https://huggingface.co/blog/AbstractPhil/beatrix-ft2&quot;&gt;https://huggingface.co/blog/AbstractPhil/beatrix-ft2&lt;/a&gt;&lt;/p&gt;
        &lt;p&gt;The control variant will need another train with better SDPA stabilization, as the control variant destabilized and collapsed. The primary fault is the lack of QK normalization, which caused the model to simply collapse given enough time. Claude lists the rest of the suspected reasons in the article.&lt;/p&gt;
        &lt;p&gt;This was a very difficult series of experiments to tune with many fault points. Trying to make heads or tails of Fable 5.1 Claude-speak hasn&#39;t been the easiest task either. It seems the model is more likely to create pedantically rigid responses rather than cooperative. Not necessarily insulting, but definitely a sort of refrigerator-magnet behavior - treating my individual contributions as little sketches for the refrigerator. This often completely ignores my larger MD or complex behavioral instructions in favor of my theoretical or hypothetical - likely considering the MD and technical as the model&#39;s own, rather than my direct contributions. Right there... right on the refrigerator goes my hypothesis that worked.&lt;/p&gt;
        &lt;p&gt;&lt;a href=&quot;https://github.com/AbstractEyes/geolip-bytelex&quot;&gt;https://github.com/AbstractEyes/geolip-bytelex&lt;/a&gt;&lt;/p&gt;
        &lt;p&gt;In any case, this upcoming week will be related entirely to cross-tokenizer distillation research. It may stretch long beyond the next week, but as it stands the geometric vocabulary has evolved into a codebook prediction system.&lt;/p&gt;
        &lt;p&gt;I would like to give this program linear wings. The Beatrix model supports it, but how well is up for this week to decide.&lt;/p&gt;
        &lt;p&gt;There are a multitude of potentials based on a series of very recent articles I will be exploring, providing the necessary bytelex complexity to a roughly 60 hour battery of experiments and trainings throughout the geometric systems.&lt;/p&gt;
        &lt;p&gt;The results will determine the best and worst methodologies of using these models, these shapes, and these structures with more complex byte-level cross tokenization systems&lt;/p&gt;
      </description>
      <link>https://huggingface.co/posts/AbstractPhil/372092347372730</link>
      <guid isPermaLink="false">372092347372730</guid>
      <pubDate>Sat, 05 Sep 2026 15:30:00 GMT</pubDate>
      <author>AbstractPhil</author>
    </item>
    <item>
      <title>Mini-Beatrix-2s pretraining is ready.</title>
      <description>&lt;p&gt;Mini-Beatrix-2s pretraining is ready.
        &lt;a href=&quot;https://huggingface.co/AbstractPhil/mini-beatrix-2s&quot;&gt;https://huggingface.co/AbstractPhil/mini-beatrix-2s&lt;/a&gt;
        The model passed a great deal of rigor and hardship, trained roughly 16 billion tokens or so. The full writeup for the model including the arms training for the first version arms and the second version arms will be drafted and prepared as soon as the v2 arms are done training and testing.&lt;/p&gt;
        &lt;p&gt;There are many possibilities present with such a model. The hub itself has been marked capable of potentially operating as similarity comparison, 87% of the capacity retained within a 256 dim structure. Along with this, the multi-dimensional hub attention system shows serious promise with controlling diffusion model inference, which I look forward to see the results of.&lt;/p&gt;
        &lt;p&gt;Additionally, sentence similarity, next token prediction, and a large array of prediction formats have been heavily improved by introducing the full model with splat attention. The model not only improved, the structure complemented everything measured, along with the more effective training regiment for version 2.&lt;/p&gt;
        &lt;p&gt;Beatrix 2s is essentially an autoregression decoder, however the attention mechanism houses a dual-stage encoder/decoder structure internally. Each adopting the SVAE as a core component, revamped and fitted to the exact rules of AlephLM. So there are essentially 20 SVAE in this structure, each with their own independent encoders, residually learning from the last.&lt;/p&gt;
        &lt;p&gt;Upcoming tests will include finetunes to bring out the strengths of all special tokens, presented in the upcoming article. The full experiment battery will be completed within a few days and the findings presented.&lt;/p&gt;
        &lt;p&gt;Modularization, compartmentalization, secularized behavior, and everything between are to be tested with rigor. This model is a rapid learner, there will likely be byproduct problems with that, and I look forward to solving the corewise problems one at a time until the model is strong enough to be useful for all the tested tasks.&lt;/p&gt;
      </description>
      <link>https://huggingface.co/posts/AbstractPhil/633488241482182</link>
      <guid isPermaLink="false">633488241482182</guid>
      <pubDate>Mon, 31 Aug 2026 20:08:42 GMT</pubDate>
      <author>AbstractPhil</author>
    </item>
  </channel>
</rss>

@github-actions github-actions Bot added the auto: ready to review Human review will come in after lint issues and merge conflicts are fixed label Sep 11, 2026

@DIYgod DIYgod left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR! One change before we merge:

  • Please switch to the JSON API instead of parsing the profile page's data-props. The server-rendered props only include the latest 2 posts, so this route can never return more than 2 items. https://huggingface.co/api/recent-activity?feedType=user&entity=<user>&activityType=post&limit=100 works without authentication and returns the user's post activity. Keep the entries with type === 'social-post' and use socialPost.rawContent, socialPost.publishedAt, socialPost.url and socialPost.slug. For AbstractPhil that currently gives 21 posts instead of 2, and you no longer need cheerio.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto: ready to review Human review will come in after lint issues and merge conflicts are fixed route

Projects

None yet

Development

Successfully merging this pull request may close these issues.

huggingface posts

3 participants