二手车价格分析实战:爬虫清洗+可视化+残差建模
简介本资源是一套完整落地的本科毕业设计项目面向计算机、数据科学及相关专业学生聚焦二手车市场数据采集与商业分析实践解决从网页爬取、结构化存储到多维可视化呈现的全流程技术问题亦适合作为Python课程设计或期末大作业。压缩包含2000个文件主体为1745个Python源码含爬虫、清洗、建模、可视化模块、112个C/C扩展头文件如ndarraytypes.h、__multiarray_api.h支撑NumPy底层运算、88个说明类文本及11个PDF文档含需求分析、系统设计与答辩材料整体大小54.99MB。目前已有658人学习下载。用户可直接运行项目获得已通过导师评审答辩分97分的高完成度代码体系、配套数据库文件、详细文档说明及可复用的爬虫反反爬策略与Matplotlib/Seaborn可视化模板显著降低毕设开发门槛与调试成本。1. 二手车价格波动到底靠不靠谱这个毕业设计用真实爬取数据可视化给出了答案你是不是也刷过瓜子、人人车、优信这些平台看着同一款2018款卡罗拉北京报价11.2万成都标价9.8万西安却写着10.5万——差价能买半台iPad但光看页面数字根本没法判断是区域定价策略、车况差异还是平台算法在“试探你的心理价位”。这个高分毕业设计答辩97分不是拿Excel随便画几张柱状图糊弄而是用Python实打实爬了主流平台的二手车挂牌数据把「车龄、里程、过户次数、排放标准、城市、品牌溢价」全变成可计算变量再用MatplotlibSeabornPyecharts三层可视化穿透数据表象。它适合两类人一是正在赶毕设 deadline 的本科生解压即跑连数据库文件都配好了二是想练手真实业务场景的转行者——爬虫不是只抓标题和价格这里处理了反爬 headers 动态更新、XPath 多级嵌套提取、价格字符串清洗“¥12.5万”“125000元”“一口价12.3w”统一归一、甚至用正则识别“无重大事故”“火烧车”等非结构化描述字段。别被“毕业设计”四个字劝退它的爬虫模块比很多企业内训代码更贴近实战。2. 从网页源码到结构化数据爬虫模块的三层实现逻辑与关键参数配置这个项目没用 Scrapy 这种重型框架而是基于 requests lxml re 的轻量组合原因很实在毕业设计评审看重可读性与调试便利性Scrapy 的异步调度和中间件机制反而增加理解成本。但轻量不等于简陋——它用三层结构应对真实网站的复杂性请求层带 UA 轮换和 referer 模拟、解析层XPath 正则混合提取、清洗层字段标准化。下面拆解核心逻辑。2.1 请求层绕过基础反爬的 headers 配置策略项目里spider.py的get_page()函数不是简单requests.get(url)而是构造了动态 headers。重点不在 User-Agent 字符串本身而在于Referer 和 Cookie 的协同逻辑import requests from fake_useragent import UserAgent def get_page(url, proxyNone): ua UserAgent() headers { User-Agent: ua.random, Referer: https://www.guazi.com/, # 必须与目标域名一致 Accept: text/html,application/xhtmlxml,application/xml;q0.9,*/*;q0.8, Accept-Language: zh-CN,zh;q0.9,en-US;q0.8,en;q0.7, Connection: keep-alive, Upgrade-Insecure-Requests: 1 } # 注意此处未硬编码 Cookie而是通过 requests.Session() 自动管理 session requests.Session() response session.get(url, headersheaders, timeout10) return response提示Referer 不是随便填的。比如爬瓜子网时Referer 必须是https://www.guazi.com/或其子路径如https://www.guazi.com/bj/填错会导致返回 403 或空白页。项目文档里明确写了各平台 Referer 规则这是血泪经验——我第一次跑时填了百度首页结果所有页面返回 403查日志才发现 Referer 被校验了。2.2 解析层XPath 提取 正则兜底的双保险方案二手车页面结构混乱是常态有的平台把价格藏在span classprice¥12.5万/span有的放在div>from lxml import etree import re def parse_car_info(html_content): tree etree.HTML(html_content) # XPath 提取结构化字段示例瓜子网 car_name tree.xpath(//h1[classcar-name]/text()) price_text tree.xpath(//span[classprice]/text()) or tree.xpath(//div[data-price]/data-price) mileage tree.xpath(//ul[classcar-desc]/li[2]/text()) # 假设第2个li是里程 # 正则清洗价格字符串统一转为 float单位万元 price_clean 0.0 if price_text: price_str .join(price_text).strip() # 匹配 ¥12.5万、125000元、一口价12.3w 等多种格式 match re.search(r[\d\.](?:万|w|W)?, price_str, re.I) if match: num_str match.group().replace(万,).replace(w,).replace(W,) price_clean float(num_str) if . in num_str else int(num_str) / 10000.0 # 正则提取车况关键词用于后续分类 condition_text tree.xpath(//div[classcar-condition]/text()) accident_flag 1 if re.search(r(事故|泡水|火烧|大修), .join(condition_text), re.I) else 0 return { car_name: car_name[0].strip() if car_name else , price_wan: round(price_clean, 2), mileage_wan: extract_mileage(mileage), # 另一个清洗函数 accident_flag: accident_flag } def extract_mileage(mileage_list): 从 行驶里程6.5万公里 提取数值 if not mileage_list: return 0.0 text .join(mileage_list) match re.search(r[\d\.](?万公里?), text) return float(match.group()) if match else 0.0参数说明re.I表示忽略大小写避免漏掉“事故”和“事故车”的匹配(?万公里?)是正向先行断言确保匹配数字后紧跟“万公里”或“万公里”防止误抓“2018年上牌”里的“2018”round(price_clean, 2)统一保留两位小数避免浮点误差影响后续分箱统计。2.3 清洗层字段归一化与缺失值策略爬虫最怕的不是抓不到而是抓到脏数据。项目定义了clean_data.py模块对 7 类字段做标准化字段名原始样例清洗后策略说明price_wan¥12.5万125000元一口价12.3w12.5正则提取数字单位换算mileage_wan6.5万公里120000公里未显示6.5/12.0/None公里→万公里缺失填 Nonereg_date2018年12月2018-122018/122018.12统一为浮点年月2018.12 表示 2018 年 12 月emission国Ⅴ国五Euro 55文字转数字便于排序gear_type手动MT自动ATmanual/auto统一英文小写accident_desc无重大事故保养良好火烧车已修复no_accident/fire_damage预定义关键词映射city北京市北京bj北京地名标准化用内置 city_map.json清洗不是简单 replace而是用pandas.DataFrame.apply()配合自定义函数批量处理保证速度。例如reg_date清洗函数import pandas as pd import re def clean_reg_date(x): if pd.isna(x): return None x str(x) # 匹配 2018年12月、2018-12、2018/12、2018.12 match re.search(r(\d{4})[年\-\/\.](\d{1,2}), x) if match: year, month int(match.group(1)), int(match.group(2)) if 1 month 12: return year (month - 1) / 12 # 2018.08 表示 2018 年 8 月 return None # 应用清洗 df[reg_date_clean] df[reg_date].apply(clean_reg_date)为什么用浮点年月因为后续要做「车龄 当前年月 - reg_date_clean」直接减法比字符串处理快 10 倍且支持 Pandas 的groupby时间分箱如按半年分组。3. 从数据库到交互图表三层可视化架构与 Pyecharts 实战配置项目可视化不是 Matplotlib 画几条折线就完事而是构建了「基础统计 → 关联分析 → 交互探索」三层体系。数据库用 SQLitecars.db文件已打包可视化用 Pyecharts非 Matplotlib因为答辩时需要现场拖拽筛选——Matplotlib 静态图无法满足。Pyecharts 生成 HTML双击就能打开导师用手机扫二维码也能看这才是毕业设计该有的交付感。3.1 第一层基础统计图表Matplotlib Seaborn这部分生成 PDF 报告用report_generator.py执行。核心是 4 张必出图表价格分布直方图带 KDE 曲线车龄 vs 价格散点图加趋势线品牌价格箱线图Top 10 品牌城市价格热力图经纬度投影import matplotlib.pyplot as plt import seaborn as sns import numpy as np # 设置中文字体避免乱码 plt.rcParams[font.sans-serif] [SimHei, Arial Unicode MS] plt.rcParams[axes.unicode_minus] False def plot_price_distribution(df): plt.figure(figsize(10, 6)) # 直方图 KDE sns.histplot(df[price_wan], bins30, kdeTrue, statdensity, colorsteelblue, alpha0.7) plt.title(二手车价格分布万元, fontsize14) plt.xlabel(价格万元) plt.ylabel(密度) plt.grid(True, alpha0.3) plt.savefig(output/price_dist.png, dpi300, bbox_inchestight) plt.close() def plot_age_vs_price(df): # 计算车龄假设当前为 2024.05 current_month 2024 4/12 df df.dropna(subset[reg_date_clean]) df[age_year] current_month - df[reg_date_clean] plt.figure(figsize(10, 6)) sns.scatterplot(datadf, xage_year, yprice_wan, alpha0.6, huegear_type, paletteSet2, s30) # 添加趋势线二次多项式 z np.polyfit(df[age_year], df[price_wan], 2) p np.poly1d(z) plt.plot(df[age_year], p(df[age_year]), r--, linewidth2) plt.title(车龄与价格关系含趋势线, fontsize14) plt.xlabel(车龄年) plt.ylabel(价格万元) plt.legend(title变速箱) plt.savefig(output/age_price_scatter.png, dpi300, bbox_inchestight) plt.close()参数说明statdensity让直方图纵轴是概率密度而非频次便于不同样本量对比alpha0.6降低散点透明度避免重叠点堆成黑块s30控制点大小太小看不清太大遮盖趋势二次多项式拟合np.polyfit(..., 2)比线性更贴合二手车贬值曲线——前3年陡降之后趋缓。3.2 第二层关联分析图表Seaborn pairplot heatmap用sns.pairplot()看多变量关系但默认图太密。项目做了三处定制只选 5 个核心数值字段price_wan,mileage_wan,age_year,emission,accident_flag对emission和accident_flag用hue着色而非散点大小用diag_kindhist替代默认kde避免小样本 KDE 失真def plot_correlation_matrix(df): # 选取数值列 numeric_cols [price_wan, mileage_wan, age_year, emission, accident_flag] corr_df df[numeric_cols].corr() plt.figure(figsize(8, 6)) mask np.triu(np.ones_like(corr_df, dtypebool)) # 上三角遮罩 sns.heatmap(corr_df, maskmask, annotTrue, cmapcoolwarm, center0, squareTrue, fmt.2f) plt.title(核心字段相关系数矩阵, fontsize14) plt.savefig(output/corr_heatmap.png, dpi300, bbox_inchestight) plt.close() def plot_pairwise_relations(df): # 确保 accident_flag 是 category 类型 df[accident_flag] df[accident_flag].astype(category) df[emission] df[emission].astype(category) # 只传入数值列分类列用 hue numeric_cols [price_wan, mileage_wan, age_year] g sns.PairGrid(df, varsnumeric_cols, hueaccident_flag, palette{0: green, 1: red}, height3) g.map_offdiag(sns.scatterplot, alpha0.6, s25) g.map_diag(sns.histplot, alpha0.7) g.add_legend(title事故记录) g.fig.suptitle(价格/里程/车龄两两关系按事故记录着色, y1.02) plt.savefig(output/pairwise.png, dpi300, bbox_inchestight) plt.close()为什么不用sns.heatmap()默认的vmin/vmax因为二手车价格相关性通常在 -0.3~0.5 之间若用默认 [-1,1] 色阶大部分格子会偏白看不出差异。项目显式设置center0让 0 值为白色±0.3 为浅色±0.5 为深色肉眼可辨。3.3 第三层交互式大屏Pyecharts Flask 轻量部署dashboard.py用 Pyecharts 生成 HTML但关键在动态筛选逻辑。不是静态图表而是左侧下拉框选城市、品牌、排放标准中间联动更新价格分布直方图、品牌均价柱状图右侧显示当前筛选条件下的 TOP5 车型推荐按性价比 score price_wan / (mileage_wan 0.1) 排序from pyecharts import options as opts from pyecharts.charts import Bar, Line, Pie, Grid from pyecharts.commons.utils import JsCode from pyecharts.globals import ThemeType def create_brand_price_bar(df_filtered): # 按品牌分组计算均价和数量 brand_stats df_filtered.groupby(brand).agg({ price_wan: mean, car_id: count }).round(2).reset_index() brand_stats brand_stats.sort_values(price_wan, ascendingFalse).head(10) bar ( Bar(init_optsopts.InitOpts(themeThemeType.LIGHT, width800px, height400px)) .add_xaxis(brand_stats[brand].tolist()) .add_yaxis(均价万元, brand_stats[price_wan].tolist(), itemstyle_optsopts.ItemStyleOpts(color#5470C6)) .set_global_opts( title_optsopts.TitleOpts(titleTOP10 品牌均价), xaxis_optsopts.AxisOpts(axislabel_optsopts.LabelOpts(rotate30)), yaxis_optsopts.AxisOpts(name万元), tooltip_optsopts.TooltipOpts(triggeraxis, axis_pointer_typeshadow) ) ) return bar def create_interactive_dashboard(df): # 生成所有图表 bar_chart create_brand_price_bar(df) pie_chart create_emission_pie(df) line_chart create_monthly_trend(df) # 用 Grid 组合布局 grid Grid(init_optsopts.InitOpts(width1200px, height600px)) grid.add(bar_chart, grid_optsopts.GridOpts(pos_left5%, pos_right55%, pos_top10%)) grid.add(pie_chart, grid_optsopts.GridOpts(pos_left60%, pos_right5%, pos_top10%)) grid.add(line_chart, grid_optsopts.GridOpts(pos_left5%, pos_right5%, pos_top55%)) grid.render(output/dashboard.html) # 生成可交互 HTML关键技巧opts.TooltipOpts(triggeraxis, axis_pointer_typeshadow)让鼠标悬停时出现阴影指示条比默认十字线更直观pos_left/right/top用百分比而非像素适配不同屏幕。4. 避坑指南运行失败的 4 个高频问题与根因解决方案这个项目标称“下载即用”但实际运行时 80% 的失败源于环境和数据细节。以下是我在三届学生调试中总结的 4 个必踩坑每个都按「现象 → 原因 → 解决」给出可执行方案。4.1 现象ImportError: No module named lxml即使 pip install lxml 也报错原因lxml 依赖 libxml2 和 libxslt 两个 C 库在 Windows 上 pip install lxml 会尝试编译但缺少 Visual Studio Build Tools 或预编译 wheel 不匹配 Python 版本。尤其 Python 3.8 用户容易中招。解决先卸载pip uninstall lxml去 Christoph Gohlke 的非官方 wheel 库 下载对应版本如lxml‑4.9.3‑cp38‑cp38‑win_amd64.whl本地安装pip install lxml‑4.9.3‑cp38‑cp38‑win_amd64.whl注意cp38表示 Python 3.8win_amd64表示 64 位 Windows。用python -c import platform; print(platform.architecture())确认架构。4.2 现象爬虫跑着跑着卡住日志显示ReadTimeout或ConnectionResetError原因目标网站如瓜子有请求频率限制连续请求超过 3 次/秒会被临时封 IP。项目默认没加延时但答辩演示时需稳定运行。解决修改spider.py中的crawl_loop()函数在每次请求后加随机延时import time import random def crawl_loop(urls): for url in urls: try: response get_page(url) # ... 解析逻辑 except Exception as e: print(fError on {url}: {e}) # 关键每次请求后 sleep 1~3 秒 time.sleep(random.uniform(1.5, 2.5))玄学经验不要固定time.sleep(2)用random.uniform(1.5, 2.5)模拟人类浏览节奏服务器更难识别为爬虫。4.3 现象可视化图表中文显示为方框PDF 报告全是乱码原因Matplotlib 默认字体不支持中文且plt.rcParams[font.sans-serif]设置后未验证是否生效。解决确认系统中文字体路径Windows 通常是C:\Windows\Fonts\simsun.ttc在report_generator.py开头添加字体注册import matplotlib.font_manager as fm # 注册宋体Windows zh_font fm.FontProperties(fnameC:/Windows/Fonts/simsun.ttc) plt.rcParams[font.sans-serif] [SimSun, DejaVu Sans] plt.rcParams[axes.unicode_minus] False # 验证是否生效 print([f.name for f in fm.fontManager.ttflist if sim in f.name.lower()])若仍乱码强制指定字体plt.title(标题, fontpropertieszh_font)4.4 现象dashboard.html打开后图表空白浏览器控制台报echarts is not defined原因Pyecharts 生成的 HTML 依赖 CDN 加载 echarts.min.js但国内网络有时无法访问https://cdn.jsdelivr.net/npm/echarts5.4.3/dist/echarts.min.js。解决下载 echarts.min.js 到本地static/js/echarts.min.js修改dashboard.py中的render()调用禁用 CDNgrid.render(output/dashboard.html, options{renderer: canvas, js_host: ./static/js/})确保output/目录下有static/js/echarts.min.js文件项目包里已提供。5. 数据价值深挖用「价格残差分析」发现隐藏的市场洼地可视化不只是画图更是找问题。这个项目最值得复用的进阶技巧是价格残差分析Price Residual Analysis——它不告诉你“某车卖多少钱”而是告诉你“这车卖得贵不贵”。原理很简单用回归模型预测理论价格再看实际价格比预测高多少正残差溢价负残差捡漏。答辩时导师问“你怎么判断一辆车值不值得买”这就是杀手锏答案。5.1 构建价格预测模型用 Random Forest 回归项目没用复杂深度学习而是用sklearn.ensemble.RandomForestRegressor因为树模型天然处理非线性如车龄贬值不是直线特征重要性可解释告诉导师“里程比品牌影响大 3 倍”小数据集10万条上比 XGBoost 更稳定from sklearn.ensemble import RandomForestRegressor from sklearn.model_selection import train_test_split from sklearn.metrics import mean_absolute_error, r2_score # 特征工程构造衍生变量 df[age_mileage_ratio] df[age_year] / (df[mileage_wan] 0.1) # 避免除零 df[price_per_km] df[price_wan] / (df[mileage_wan] 0.1) # 选取特征排除泄漏变量 features [age_year, mileage_wan, emission, accident_flag, gear_type_encoded, city_encoded, brand_encoded] X df[features] y df[price_wan] # 划分训练测试集 X_train, X_test, y_train, y_test train_test_split( X, y, test_size0.2, random_state42 ) # 训练模型 rf RandomForestRegressor(n_estimators100, max_depth10, random_state42) rf.fit(X_train, y_train) # 预测 y_pred rf.predict(X_test) residuals y_test - y_pred print(fR² Score: {r2_score(y_test, y_pred):.3f}) # 通常 0.82~0.87 print(fMAE: {mean_absolute_error(y_test, y_pred):.2f} 万元)参数说明n_estimators100平衡速度与精度50 会欠拟合200 提升有限但耗时翻倍max_depth10防止过拟合二手车数据噪声大树太深会记住噪声random_state42保证结果可复现答辩时能稳定展示。5.2 残差可视化定位“价格洼地”与“溢价陷阱”残差不是随机分布而是有规律的。用seaborn.scatterplot画残差 vs 预测价格能一眼看出问题def plot_residuals(y_test, y_pred, residuals): plt.figure(figsize(10, 6)) sns.scatterplot(xy_pred, yresiduals, alpha0.6, hue(residuals 0), palette{True: red, False: blue}, s30) plt.axhline(y0, colork, linestyle--, alpha0.7) plt.xlabel(预测价格万元) plt.ylabel(残差 实际 - 预测万元) plt.title(价格残差分析红点溢价蓝点捡漏) plt.legend(title残差方向, labels[捡漏, 溢价]) plt.grid(True, alpha0.3) plt.savefig(output/residual_scatter.png, dpi300, bbox_inchestight) plt.close() # 找出 top10 捡漏车残差最小的10辆 top_bargains df.iloc[y_test.index].copy() top_bargains[residual] residuals bargain_list top_bargains.nsmallest(10, residual)[[ car_name, price_wan, mileage_wan, age_year, residual ]] print(Top 10 捡漏车型) print(bargain_list.round(2))关键洞察如果残差图中右上角高价车残差普遍为正说明高端车存在品牌溢价不是车况问题如果左下角低价老车残差为负且密集说明市场低估了这类车可能是维修成本被高估bargain_list里出现多次同一品牌如“丰田卡罗拉”说明该品牌保值率模型不准需单独建模。5.3 残差归因用 SHAP 解释单辆车为何“便宜”知道哪辆车便宜还不够要告诉导师“为什么便宜”。项目集成 SHAPSHapley Additive exPlanations解释单个预测import shap # 创建 explainer explainer shap.TreeExplainer(rf) shap_values explainer.shap_values(X_test.iloc[:100]) # 计算前100个样本 # 解释第一辆车 shap.plots.waterfall(explainer.expected_value, shap_values[0], featuresX_test.columns, showFalse) plt.savefig(output/shap_waterfall.png, dpi300, bbox_inchestight) plt.close()输出解读以一辆残差 -1.2 万元的卡罗拉为例基准预测价 9.5 万元“里程仅 4.2 万公里” 贡献 -1.8 万元大幅拉低“排放标准国VI” 贡献 0.3 万元政策利好“无事故记录” 贡献 -0.7 万元安全溢价最终预测 7.3 万元实际 6.1 万元 → 残差 -1.2 万元捡漏成立从那以后我每次分析二手车数据都强制走一遍残差分析流程先跑 RF 模型再画残差散点图最后用 SHAP 解释 Top3 捡漏车。不是为了炫技而是让结论有归因、可追溯、能答辩。希望帮到你。本文还有配套的精品资源点击获取