Piotr VassevPiotr Vassev•

How to Scrape Weibo Posts, Comments, and Profiles

A Weibo post's like and comment counts tell you it got attention. They don't tell you who responded, what they said, or where they posted from. For that you need the comments themselves, next to the post and its author.

Our Weibo Scraper exports all three without a Weibo account. Give it post links and it returns the post, its public comments with the first replies, and profiles for any users you list. It can also export the live hot-search board (热搜). It cannot search Weibo by keyword or list an account's full post history, because Weibo keeps both behind a login.

Weibo posts, comments, profiles, and the hot-search board exported to a dataset

Pick a mode

The Actor runs one mode at a time, chosen with the What to scrape field (mode in JSON):

ModeInputOne row per
postDetailsPost links in postUrlsPost, with long posts expanded to full text
postCommentsPost links in postUrlsTop-level comment, with its first replies inside the row
userProfilesUsers in usersAccount
hotSearchNothingHot-search topic
hotTimeline (default)NothingPost from Weibo's hot feed

Post links can be weibo.com/{user}/{id}, m.weibo.cn/detail/{id}, m.weibo.cn/status/{id}, or a bare post id. Users can be weibo.com/u/{uid} links, custom links such as weibo.com/rmrb, numeric ids, or screen names like 人民日报.

Every row carries a type field (post, comment, profile, hotSearch), so rows stay identifiable after you merge exports.

Budget before you run

Pricing as of September 28, 2026: $0.005 per dataset row, whatever the row type. A post, a comment, a profile, and a hot-search topic each count as one row. The three runs below (1 post, up to 100 comments, 1 profile) cost at most 102 × $0.005 = $0.51. Check current pricing before scheduling anything recurring.

Two limits control the cost. limit caps the whole run (default 100). maxCommentsPerPost caps comments on each post (default 100). With three posts and maxCommentsPerPost: 100, you also need limit of at least 300 to get all of them.

Export one post's reaction

We'll use a post from the actor 曾舜晞 about the Greater Bay Area autumn gala (大湾区秋晚). Our example run showed 430,206 likes and 139,386 comments on it.

Open the Actor on Apify, go to Input, switch to the JSON editor, and run these three inputs one after another.

1. The post:

{
  "mode": "postDetails",
  "postUrls": ["https://weibo.com/1763990660/RjANMzWLc"]
}

2. Its comments:

{
  "mode": "postComments",
  "postUrls": ["https://weibo.com/1763990660/RjANMzWLc"],
  "maxCommentsPerPost": 100
}

3. The author's profile:

{
  "mode": "userProfiles",
  "users": ["https://weibo.com/u/1763990660"]
}

Leave Proxy configuration at its default. After each run finishes, open the run's Dataset and export JSON. JSON keeps Chinese text and the nested author and replies fields intact. If you prefer CSV, open it as UTF-8 and spot-check a few Chinese names.

The run log ends each post with a line like Post RjANMzWLc: 100 comments (post shows 139386). Compare the two numbers before you draw conclusions. The next section explains why they differ.

Read the rows

A comment row, abbreviated from our example run:

{
  "type": "comment",
  "postId": "5346697028567358",
  "postUrl": "https://weibo.com/1763990660/RjANMzWLc",
  "commentId": "5346697106425782",
  "text": "秋晚想见,和曾舜晞",
  "createdAt": "Thu Sep 24 15:36:30 +0800 2026",
  "likes": 1966,
  "location": "浙江",
  "author": {
    "id": "7884193234",
    "name": "今天也要为曾舜晞拼命",
    "verified": true,
    "followersCount": 13482
  },
  "replyCount": 3,
  "replies": [
    { "commentId": "5346697141027365", "text": "相见…", "likes": 905, "location": "浙江" }
  ]
}

location is the province Weibo displays under the comment ("来自 浙江"), with the prefix removed. replies holds the first few replies Weibo shows with the comment, not the full reply thread. replyCount is the thread's total.

A profile row gives the numbers post rows lack. For 人民日报 our example run returned:

FieldValue
followersCount158,011,374
postsCount153,651
totalLikes2,527,284,539
verifiedReason《人民日报》法人微博
registeredAt2012-07-22 02:28:35
credit阳光信用极好

company and education fill in only when the account lists them.

Why the comment count won't match

Weibo shows logged-out visitors part of each thread, not all of it. The Actor walks the thread newest-first, then fills in from Weibo's hot order, dropping duplicates. In our example runs that returned 182 of 217 comments on a smaller post and 400 of 400 requested on the 139,386-comment post above. A request for 5,000 comments on a post of that size will stop well short.

Treat the export as a large sample of the visible discussion, not a census. Write the collection date and row count next to any chart you make from it.

One more case: an author can turn on comment curation (评论精选), which publishes only the comments they pick. We saw a post with 243 comments return one. The run log adds "author shows selected comments only" when that happens.

Build the reaction table

Join the three exports on their ids: comment postId to post postId, and post author.id to profile userId. Keep all ids as text in spreadsheets, since 16-digit ids get rounded as numbers.

From there, a few checks are worth doing before any sentiment scoring:

  • Where replies come from. Count comments by location. A thread dominated by one or two provinces may reflect a fan base more than general opinion.
  • Who is replying. Filter author.verified and sort by author.followersCount to see which established accounts joined in, then run userProfiles on the ones that matter.
  • What rose to the top. Sort by likes. The most-liked comments set the tone most readers see. With a small maxCommentsPerPost, add "commentsSort": "hot" to the comments input: the default newest-first order starts with fresh comments that have few likes yet.
  • Reach versus response. Compare the post's likes and comments with the author's followersCount from the profile row.

Read the original Chinese before labeling tone. Machine translation often loses sarcasm, fandom slang, and wordplay.

Take a hot-search snapshot

For a view of what Weibo is discussing right now, run:

{
  "mode": "hotSearch"
}

Our example run returned 51 rows: one pinned topic with pinned: true and no heat value, then 50 ranked topics with rank, keyword, heat, and label (such as 新 for new or 热 for hot). Each row has a searchUrl that opens that topic on Weibo in a browser. At current pricing a full board costs about $0.26.

The board changes throughout the day. To track a topic's rise, schedule the run in Apify and compare rank and heat across snapshots by keyword. The Actor doesn't collect posts for a topic, so pick posts from the topic page yourself and feed their links to postComments.

Weibo hot timeline output with post text, engagement counts, and author fields

When a run returns less than expected

  • Nothing at all: check that the proxy is still enabled. Weibo refuses Apify's platform IPs without one.
  • "Post … was not found": the post was deleted, hidden, or restricted to logged-in users. The run skips it and continues.
  • A user not found: try the numeric id from their weibo.com/u/ link instead of the screen name.
  • followersCount: null on post rows: expected. Run userProfiles for those authors.
  • Broken image or video links: media URLs carry expiry parameters. Keep postUrl as the lasting reference.

For running the Actor from code, see the Node.js example. It runs the three steps above for any post link, prints comments by location and the most-liked ones, and saves a hot-search snapshot with npm run hot.

Frequently asked questions

Can I search Weibo by keyword or export an account's full post history?

No. Weibo only serves keyword search and full user timelines to logged-in accounts, and our Actor does not log in. Start from post links, user profiles, the hot-search board, or the hot timeline instead.

Why do I get fewer comments than the post shows?

Weibo shows logged-out visitors only part of a comment thread. Our example runs returned 182 of 217 comments on a small post and 400 of 400 requested on a post showing about 139,000. Some authors also turn on comment curation (评论精选), which leaves only the comments they pick visible.

Why is followersCount null on post rows?

Weibo's feed and post payloads do not include the author's follower count. Run the userProfiles mode for those authors to get followersCount, following, post counts, and lifetime interaction totals.

Do I need a proxy or a Weibo account?

You need a proxy but no account. Weibo does not answer Apify's shared platform IPs, so keep the default Apify Proxy setting. No cookies or login are required.

Piotr Vassev

Piotr Vassev

Founder of FalconScrape. Building production-grade web scraping systems and data automation pipelines for businesses worldwide.

Connect on LinkedIn