# 社会主义经典著作集 · 接口文档

面向第三方开发者的公开接口说明。

**站点**：https://rmws1976.xyz
**数据**：18,596 篇文章 · 14 位作者 · 正文合计 66,348,218 字

---

## 一、接口一览

一共 4 个接口，全部为 **GET**：

| 接口 | 作用 | 特点 |
|---|---|---|
| `/search` | **语义检索** —— 找"意思相近"的文章 | 向量相似度 + 词法加权；结果按相关性排序 |
| `/text` | **全文检索** —— 找"哪篇文章写了这个词" | 字面匹配；返回每篇的**出现次数**和上下文片段 |
| `/article/{id}` | 取一篇文章的**全文** | 轻量，不排队 |
| `/author/{name}` | 取某作者的**文章目录** | 轻量，不排队；支持按日期筛 |

文章 `id` 在全站统一：`/search`、`/text`、`/author` 返回的 `id`，与网页地址 `/p/{id}.html` 一一对应。

### 语义检索还是全文检索

这两个接口**互补，不是替代**：

- 问「关于 X 的论述」「X 的思想」→ 用 `/search`
- 问「哪篇文章里出现过『亚罗号』这个词」「出现几次」→ 用 `/text`

同一条问法在一个接口下没有结果、在另一个接口下命中很多，是常态。

---

## 二、通用约定

### 请求

- 全部为 GET，参数放在 query string。
- 中文参数需 URL 编码（标准 `urlencode` 即可，服务端也做了容错）。
- 服务端未开启 CORS，**仅供服务端与脚本调用**，不适合网页前端直接跨域请求。

### 返回里的判断字段

每个接口都会返回 `verdict`，用来判断查到了没有：

| 值 | 含义 |
|---|---|
| `hit` | 有命中 |
| `weak` | 命中但相关度偏弱 |
| `absent` | 没有命中 |

> ⚠️ **「没有命中」是 HTTP 200 + `verdict: "absent"`，不是 404。**
> 请用 `verdict` 判断，不要用 HTTP 状态码。

### limit

默认值：`/search` 为 200，`/text` 与 `/author` 为 50。上限 **1000**。

`/search` 支持 `limit=0`：不返回内容，只返回准确的 `total`，适合先探数量。

### 排队与 503

`/search` 与 `/text` 是 CPU 密集操作，服务端做了并发限制。请求过多时会排队，**排队超过 30 秒**返回：

```json
HTTP 503
{"verdict": "error", "busy": true,
 "message": "服务繁忙：排队超过 30 秒仍未轮到，请稍后重试"}
```

> ⚠️ **请把 503 当作「服务繁忙」重试，不要当作「没有内容」。**
> 判断依据是 `busy: true`。

`/article` 与 `/author` 不排队。

### 响应头

| 头 | 说明 |
|---|---|
| `X-Queue-Wait-Ms` | 本次请求排队等待的毫秒数 |
| `X-Inflight` | 返回时正在处理的请求数 |

---

## 三、`GET /search` —— 语义检索

### 参数

| 参数 | 类型 | 默认 | 说明 |
|---|---|---|---|
| `q` | string | `''` | 检索词。**可以为空** —— 空 `q` 时退化为"按作者/范围列出文章"，按日期升序 |
| `author` | string | `''` | 按作者精确过滤 |
| `authors` | string | `''` | 多作者，逗号分隔，如 `马克思,恩格斯`。`author` 优先级更高 |
| `scope` | string | `''` | 只认 `classics`（经典作家范围）。见下 |
| `date_from` | string | `''` | 起始日期，`YYYY-MM-DD` |
| `date_to` | string | `''` | 结束日期，`YYYY-MM-DD` |
| `offset` | int | 0 | 翻页偏移，见「翻页」 |
| `limit` | int | 200 | 返回条数（上限 1000）；`limit=0` 只返回 `total` |

#### `scope` 说明

`scope` **只在没有传 `author`/`authors` 时才可能生效**，且只认 `classics`。

传了其它值不会报错，但会在返回里给出 `scope_ignored` 字段说明它被忽略了 —— 请检查该字段。

#### 周群

周群的文章无法确定具体年份，查询该作者时服务端会**忽略日期过滤**，并在返回里给出 `date_ignored` / `date_note`。

### 返回

```json
{
  "query": "鸦片战争",
  "author": "马克思",
  "authors": ["马克思"],
  "scope": "author",
  "verdict": "hit",
  "total": 12,
  "count": 2,
  "offset": 0,
  "next_offset": 2,
  "top_sim": 0.6584,
  "weak_floor": 0.52,
  "lex_hits": 3,
  "sub_hits": 0,
  "results": [
    {
      "id": 9259,
      "title": "中国革命和欧洲革命",
      "author": "马克思",
      "date": "1853-05-20",
      "url": "/p/9259.html",
      "summary": "马克思运用对立统一规律，剖析鸦片战争后中国革命对英国经济及欧洲政治的连锁影响……",
      "tags": "马克思,对立统一规律,鸦片战争,中国革命,英国经济,欧洲政治……",
      "sim": 0.5847,
      "match": "tag"
    }
  ]
}
```

| 字段 | 说明 |
|---|---|
| `total` | 命中的文章总数（不受 `limit` 影响） |
| `count` | 本次返回的条数 |
| `offset` / `next_offset` | 翻页用，见下 |
| `verdict` | `hit` / `weak` / `absent` |
| `top_sim` | 最高相似度；**空 `q` 的列表模式下为 `null`** |
| `weak_floor` | 弱相关判定下限，固定 `0.52` |
| `lex_hits` / `sub_hits` | 词法/子串加权命中的文章数 |
| `results[].sim` | 该篇的相似度（0~1） |
| `results[].match` | 命中方式：`tag` / `tag_sub` / `vector`。**空 `q` 的列表模式下为布尔 `false`** |
| `results[].tags` | 关键词标签，逗号分隔，可当作极简摘要使用 |
| `scope_ignored` | 仅当传入的 `scope` 被忽略时出现 |
| `date_ignored` / `date_note` | 仅周群等日期不可靠的作者出现 |

### 翻页

用 `offset` + `next_offset`：

```
第 1 次：/search?q=中国&author=马克思&limit=5&offset=0
         → results 5 条，next_offset = 5
第 2 次：/search?q=中国&author=马克思&limit=5&offset=5
         → results 5 条，next_offset = 10
第 3 次：/search?q=中国&author=马克思&limit=5&offset=10
         → results 5 条，next_offset = null   ← 没有下一页了
```

`next_offset` 为 `null` 表示已取完，不要继续翻。

> 另有遗留参数 `before_id`（配 `next_cursor`），按文章 id 阈值翻页。
> 因为结果按相关性排序而非按 id 排序，**该方式会漏结果**，请勿使用。

### 示例

```bash
# 基本检索
curl 'https://rmws1976.xyz/search?q=%E9%B8%A6%E7%89%87%E6%88%98%E4%BA%89&author=%E9%A9%AC%E5%85%8B%E6%80%9D&limit=3'

# 只取数量
curl 'https://rmws1976.xyz/search?q=%E4%B8%AD%E5%9B%BD&author=%E9%A9%AC%E5%85%8B%E6%80%9D&limit=0'

# 按日期范围
curl 'https://rmws1976.xyz/search?q=%E4%B8%AD%E5%9B%BD&author=%E9%A9%AC%E5%85%8B%E6%80%9D&date_from=1850-01-01&date_to=1860-12-31'
```

---

## 四、`GET /text` —— 全文检索

在正文中查找**字面出现**的词，返回每篇的出现次数与上下文片段。

### 参数

| 参数 | 类型 | 默认 | 说明 |
|---|---|---|---|
| `q` | string | — | **必填**。为空时返回 `verdict: absent` 与 `detail: "q 必填"` |
| `author` | string | `''` | 按作者过滤。**建议尽量带上**，见下 |
| `limit` | int | 50 | 返回篇数（上限 1000） |

#### 性能

带 `author` 时只扫描该作者的正文，不带则扫描全部 6,634 万字，**耗时约为 8 倍**：

| 场景 | `scanned_chars` | `took_ms` |
|---|---|---|
| 单作者 | 约 546 万 | 约 5~8 ms |
| 全库 | 66,348,218 | 约 45~80 ms |

不带 `author` 时，返回中会给出 `suggest_author: true` 与 `note` 提示。

### 返回

```json
{
  "query": "亚罗号",
  "author": "马克思",
  "mode": "text",
  "verdict": "hit",
  "total": 4,
  "count": 2,
  "hits_total": 14,
  "scanned_chars": 5462879,
  "took_ms": 5.5,
  "results": [
    {
      "id": 6971,
      "title": "议会关于对华军事行动的辩论",
      "author": "马克思",
      "date": "1857-02-27",
      "url": "/p/6971.html",
      "hits": 8,
      "match": "text",
      "snippet": "……［注：克兰沃斯。——编者注］说过：“如果英国在‘亚罗号’事件上没有充分的根据，那末英国的一切行动自始至终都是错误的。”……",
      "tags": "马克思,帕麦斯顿政府,侵华战争……"
    }
  ]
}
```

| 字段 | 说明 |
|---|---|
| `total` | 命中的**文章数** |
| `hits_total` | 命中的**总次数**（同一篇多次出现累计） |
| `results[].hits` | 该篇内的出现次数 |
| `results[].snippet` | 命中处的**上下文片段**，不是全文；要全文请用 `/article/{id}` |
| `scanned_chars` / `took_ms` | 本次扫描字数与服务端耗时 |
| `note` / `suggest_author` | 仅在不带 `author` 时出现 |

---

## 五、`GET /article/{id}` —— 文章全文

| 参数 | 说明 |
|---|---|
| `id` | 文章 id（路径参数） |

```json
{
  "id": 7038,
  "title": "鸦片贸易史",
  "author": "马克思",
  "date": "1858-09-03",
  "url": "/p/7038.html",
  "tags": "英国政府,印度鸦片,中国,垄断,走私,财政利益……",
  "content": "卡·马克思\n\n正因为英国政府把在印度种植鸦片的垄断权据为己有，中国才采取了禁止鸦片贸易的措施。……"
}
```

`id` 不存在时返回 **HTTP 404**：

```json
{"detail": "not found: 99999999"}
```

> `content` 是整篇正文，长文可达两万余字，批量获取时注意体积。

---

## 六、`GET /author/{name}` —— 作者文章目录

| 参数 | 类型 | 默认 | 说明 |
|---|---|---|---|
| `name` | string | — | 作者名（路径参数，需 URL 编码） |
| `limit` | int | 50 | 返回条数 |
| `offset` | int | 0 | 偏移，翻页用 |
| `date_from` / `date_to` | string | `''` | 日期范围 |

```json
{
  "author": "马克思",
  "verdict": "hit",
  "total": 2551,
  "limit": 2,
  "offset": 0,
  "articles": [
    {"id": 1,    "title": "青年在选择职业时的考虑", "date": "1835-08",    "url": "/p/1.html"},
    {"id": 7327, "title": "根据《约翰福音》第15章……", "date": "1835-08-10", "url": "/p/7327.html"}
  ]
}
```

> ⚠️ 本接口返回的字段名是 **`articles`**（不是 `results`），每项只有
> `id` / `title` / `date` / `url` 四个字段 —— 没有摘要与标签，需要详情请另调 `/article/{id}`。

作者不存在时返回 `verdict: "absent"` 与空的 `articles`（HTTP 仍为 200）。

**已知作者（14 位）**：马克思、恩格斯、马克思/恩格斯合著、列宁、斯大林、毛泽东、胡志明、金日成、金正日、金正恩、菲德尔·卡斯特罗、周群、鲁迅、波尔布特。

---

## 七、错误与边界速查

| 情况 | HTTP | 返回 |
|---|---|---|
| 正常命中 | 200 | `verdict: "hit"` 或 `"weak"` |
| 没有命中 | 200 | `verdict: "absent"`, `total: 0` |
| `/text` 缺少 `q` | 200 | `verdict: "absent"`, `detail: "q 必填"` |
| 排队超过 30 秒 | **503** | `{"busy": true, "message": "服务繁忙：…"}` |
| `/article/{id}` 不存在 | **404** | `{"detail": "not found: …"}` |
| `/author/{name}` 不存在 | 200 | `verdict: "absent"`, `articles: []` |
| 传入不支持的 `scope` | 200 | 正常返回，另含 `scope_ignored` |

**两点最容易错**：

1. **「没有命中」是 200 而不是 404** —— 用 `verdict` 判断。
2. **503 是「繁忙」而不是「没有」** —— 用 `busy` 判断，然后重试。

---

## 八、调用示例

### Python

```python
import json, urllib.error, urllib.parse, urllib.request

BASE = "https://rmws1976.xyz"

def kb(path, **params):
    url = BASE + path + ("?" + urllib.parse.urlencode(params) if params else "")
    try:
        with urllib.request.urlopen(url, timeout=60) as r:
            return json.loads(r.read().decode("utf-8"))
    except urllib.error.HTTPError as e:
        d = json.loads(e.read().decode("utf-8", "replace"))
        d["_http"] = e.code
        return d

# 语义检索
d = kb("/search", q="鸦片战争", author="马克思", limit=3)
if d.get("busy"):
    print("服务繁忙，稍后重试")
elif d["verdict"] == "absent":
    print("没有相关内容")
else:
    print("共 %d 篇" % d["total"])
    for it in d["results"]:
        print(" ", it["id"], it["title"], "相似度 %.3f" % it["sim"])

# 全文检索（带词频）
t = kb("/text", q="亚罗号", author="马克思", limit=5)
for it in t["results"]:
    print(" ", it["title"], "出现", it["hits"], "次")

# 翻页（用 offset / next_offset）
seen, offset = [], 0
while True:
    d = kb("/search", q="中国", author="马克思", limit=50, offset=offset)
    if d.get("busy"):
        continue
    seen += [it["id"] for it in d["results"]]
    offset = d.get("next_offset")
    if offset is None:
        break
print("共取到", len(seen), "篇")

# 作者目录 → 读全文
a = kb("/author/" + urllib.parse.quote("马克思"), limit=3)
full = kb("/article/%d" % a["articles"][0]["id"])
print(full["title"], "正文", len(full["content"]), "字")
```

### Lua

```lua
-- 需要支持 HTTPS 的 HTTP 库（如 luasocket + luasec），以及一个 JSON 库
local http  = require("socket.http")
local ltn12 = require("ltn12")
local json  = require("cjson")

local function kb(path)
  local body = {}
  http.request{
    url  = "https://rmws1976.xyz" .. path,
    sink = ltn12.sink.table(body),
  }
  return json.decode(table.concat(body))
end

local d = kb("/search?q=" .. urlencode("鸦片战争") .. "&limit=5")
if d.busy then
  print("服务繁忙，稍后重试")
elseif d.verdict == "absent" then
  print("没有相关内容")
else
  print("共 " .. d.total .. " 篇")
  for i, it in ipairs(d.results) do
    print(i .. ". " .. it.title .. "  " .. it.date)
  end
end
```

> URL 里的中文需先做百分号编码。Lua 标准库没有这个函数，请用运行环境提供的方法。

### curl

```bash
curl 'https://rmws1976.xyz/search?q=%E4%B8%AD%E5%9B%BD&author=%E9%A9%AC%E5%85%8B%E6%80%9D&limit=5'
curl 'https://rmws1976.xyz/text?q=%E4%BA%9A%E7%BD%97%E5%8F%B7&author=%E9%A9%AC%E5%85%8B%E6%80%9D&limit=5'
curl 'https://rmws1976.xyz/article/7038'
curl 'https://rmws1976.xyz/author/%E9%A9%AC%E5%85%8B%E6%80%9D?limit=10'
```

---

## 附：字段速查

```
/search   query author authors scope verdict total count offset next_offset
          top_sim weak_floor lex_hits sub_hits results[]
          [scope_ignored] [date_ignored] [date_note]
results[] id title author date url summary tags sim match

/text     query author mode verdict total count hits_total
          scanned_chars took_ms results[] [note] [suggest_author]
results[] id title author date url hits match snippet tags

/article  id title author date url content tags

/author   author verdict total limit offset articles[]
articles[] id title date url
```
