Mercurial > hg > egg-tcls
annotate urllog.tcl @ 225:cb86368b8fcd
urllog: Change handling of HTTP requests.
author | Matti Hamalainen <ccr@tnsp.org> |
---|---|
date | Mon, 01 Dec 2014 11:11:50 +0200 |
parents | aaf433ab696a |
children | d0fb1546d73f |
rev | line source |
---|---|
0 | 1 ########################################################################## |
2 # | |
222 | 3 # URLLog v2.3.0 by Matti 'ccr' Hamalainen <ccr@tnsp.org> |
176
eda776bcb7ed
urllog: Bump copyright, and version.
Matti Hamalainen <ccr@tnsp.org>
parents:
160
diff
changeset
|
4 # (C) Copyright 2000-2014 Tecnic Software productions (TNSP) |
0 | 5 # |
113
077c7383f36f
urllog: Add line about the script's license.
Matti Hamalainen <ccr@tnsp.org>
parents:
112
diff
changeset
|
6 # This script is freely distributable under GNU GPL (version 2) license. |
077c7383f36f
urllog: Add line about the script's license.
Matti Hamalainen <ccr@tnsp.org>
parents:
112
diff
changeset
|
7 # |
0 | 8 ########################################################################## |
9 # | |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
10 # URL-logger script for EggDrop IRC robot, utilizing SQLite3 database |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
11 # This script requires SQLite TCL extension. Under Debian, you need: |
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
12 # tcl8.5 libsqlite3-tcl (and eggdrop eggdrop-data, of course) |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
13 # |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
14 # NOTICE! If you are upgrading to URLLog v2.0+ from any 1.x version, you |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
15 # may want to run a conversion script against your URL-database file, |
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
16 # if you wish to preserve the old data. |
0 | 17 # |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
18 # See convert_urllog_db.tcl for more information. |
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
19 # |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
20 # If you are doing a fresh install, you will need to create the |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
21 # initial SQLite3 database with the required table schemas. You |
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
22 # can do that by running: create_urllog_db.tcl |
0 | 23 # |
24 ########################################################################## | |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
25 |
0 | 26 ### |
27 ### HTTP options | |
28 ### | |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
29 # Set to 1 if you want to enable use of HTTP proxy. |
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
30 # If you do, you MUST set the proxy settings below too. |
0 | 31 set http_proxy 0 |
32 | |
33 # Proxy host and port number (only used if enabled above) | |
34 set http_proxy_host "" | |
35 set http_proxy_port 8080 | |
36 | |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
37 # Enable _experimental_ TLS/SSL support. This may not work at all. |
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
38 # If unsure, leave this option disabled (0). |
104
da337ca10e0a
urllog: Enable SSL/TLS support by default.
Matti Hamalainen <ccr@tnsp.org>
parents:
103
diff
changeset
|
39 set http_tls_support 1 |
0 | 40 |
89
77e05ce9e9b8
urllog: Add certdir option setting.
Matti Hamalainen <ccr@tnsp.org>
parents:
87
diff
changeset
|
41 set http_tls_cadir "/usr/share/ca-certificates/mozilla" |
77e05ce9e9b8
urllog: Add certdir option setting.
Matti Hamalainen <ccr@tnsp.org>
parents:
87
diff
changeset
|
42 |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
43 |
0 | 44 ### |
45 ### General options | |
46 ### | |
47 | |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
48 # Channels where URLLog records links/URLs |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
49 # set urllog_log_channels "#foobar;#baz" |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
50 # You can use * to match substrings or everything |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
51 set urllog_log_channels "*" |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
52 |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
53 |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
54 |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
55 # Filename of the SQLite URL database file |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
56 set urllog_db_file "urllog.sqlite" |
0 | 57 |
58 | |
59 # 1 = Verbose: Say messages when URL is OK, bad, etc. | |
60 # 0 = Quiet : Be quiet (only speak if asked with !urlfind, etc) | |
61 set urllog_verbose 1 | |
62 | |
63 | |
50
f69363fc1f61
Update some comments and add a bit of documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
49
diff
changeset
|
64 # 1 = Enable logging of various script actions into bot's log |
0 | 65 # 0 = Don't. |
66 set urllog_logmsg 1 | |
67 | |
68 | |
69 # 1 = Check URLs for validity and existence before adding. | |
70 # 0 = No checks. Add _anything_ that looks like an URL to the database. | |
71 set urllog_check 1 | |
72 | |
73 | |
74 ### | |
75 ### Search related settings | |
76 ### | |
77 | |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
78 # Channels where !urlfind and other commands can be used. |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
79 # By default this is set to be the same as urllog_log_channels |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
80 set urllog_search_channels $urllog_log_channels |
0 | 81 |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
82 # Limit how many URLs should the "!urlfind" command show at most. |
0 | 83 set urllog_showmax_pub 3 |
84 | |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
85 # Same as above, but for private message search. |
0 | 86 set urllog_showmax_priv 6 |
87 | |
88 | |
89 ### | |
90 ### ShortURL-settings | |
91 ### | |
181
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
92 # To enable ShortURL functionality, you need to set up the |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
93 # URL redirector PHP script (urlredirect.php) correctly, and |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
94 # enable change the settings in it and below appropriately. |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
95 # See urlredirect.php.txt for more information. |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
96 # |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
97 # You will also need SQLite3 support for PHP and access to |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
98 # change .htaccess file(s) on your web server. The PHP |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
99 # script will also need access to the SQLite3 database this |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
100 # script uses. |
ff23ce8b938f
urllog: Document ShortURL functionality slightly.
Matti Hamalainen <ccr@tnsp.org>
parents:
176
diff
changeset
|
101 # |
0 | 102 |
73
646b2fd67312
urllog: Improve documentation of different settings.
Matti Hamalainen <ccr@tnsp.org>
parents:
70
diff
changeset
|
103 # 1 = Enable showing of ShortURLs |
646b2fd67312
urllog: Improve documentation of different settings.
Matti Hamalainen <ccr@tnsp.org>
parents:
70
diff
changeset
|
104 # 0 = ShortURLs not shown in any bot actions |
0 | 105 set urllog_shorturl 1 |
106 | |
73
646b2fd67312
urllog: Improve documentation of different settings.
Matti Hamalainen <ccr@tnsp.org>
parents:
70
diff
changeset
|
107 # Max length of original URL to be shown, rest is chopped |
646b2fd67312
urllog: Improve documentation of different settings.
Matti Hamalainen <ccr@tnsp.org>
parents:
70
diff
changeset
|
108 # off if the URL is longer than the specified amount. |
0 | 109 set urllog_shorturl_orig 30 |
110 | |
73
646b2fd67312
urllog: Improve documentation of different settings.
Matti Hamalainen <ccr@tnsp.org>
parents:
70
diff
changeset
|
111 # Web server URL that handles redirects of ShortURLs |
0 | 112 set urllog_shorturl_prefix "http://tnsp.org/u/" |
113 | |
114 | |
115 ### | |
81
17e542b7985a
urllog, quotedb: Improve documentation.
Matti Hamalainen <ccr@tnsp.org>
parents:
73
diff
changeset
|
116 ### Message texts (informal, errors, etc.) |
0 | 117 ### |
118 | |
119 # No such host was found | |
120 set urlmsg_nosuchhost "ei tommosta oo!" | |
121 | |
122 # Could not connect host (I/O errors etc) | |
123 set urlmsg_ioerror "kraak, virhe yhdynnässä." | |
124 | |
125 # HTTP timeout | |
126 set urlmsg_timeout "ei jaksa ootella" | |
127 | |
128 # No such document was found | |
129 set urlmsg_errorgettingdoc "siitosvirhe" | |
130 | |
131 # URL was already known (was in database) | |
132 set urlmsg_alreadyknown "wanha!" | |
133 #set urlmsg_alreadyknown "Empiiristen havaintojen perusteella ja tällä sovellutusalueella esiintyneisiin aikaisempiin kontekstuaalisiin ilmaisuihin viitaten uskallan todeta, että sovellukseen ilmoittamasi tietoverkko-osoite oli kronologisti ajatellen varsin postpresentuaalisesti sopimaton ja ennestään hyvin tunnettu." | |
134 | |
135 # No match was found when searched with !urlfind or other command | |
136 set urlmsg_nomatch "Ei osumia." | |
137 | |
138 | |
139 ### | |
140 ### Things that you usually don't need to touch ... | |
141 ### | |
142 | |
143 # What IRC "command" should we use to send messages: | |
144 # (Valid alternatives are "PRIVMSG" and "NOTICE") | |
145 set urllog_preferredmsg "PRIVMSG" | |
146 | |
147 # The valid known Top Level Domains (TLDs), but not the country code TLDs | |
148 # (Now includes the new IANA published TLDs) | |
90
a9a4456eb213
urllog: Add .xxx TLD to supported list.
Matti Hamalainen <ccr@tnsp.org>
parents:
89
diff
changeset
|
149 set urllog_tlds "org,com,net,mil,gov,biz,edu,coop,aero,info,museum,name,pro,int,xxx" |
0 | 150 |
151 | |
152 ########################################################################## | |
153 # No need to look below this line | |
154 ########################################################################## | |
155 set urllog_name "URLLog" | |
222 | 156 set urllog_version "2.3.0" |
0 | 157 |
158 set urllog_tlds [split $urllog_tlds ","] | |
159 set urllog_httprep [split "\@|%40|{|%7B|}|%7D|\[|%5B|\]|%5D" "|"] | |
160 | |
102
5425dc418505
urllog: Entity data is now in UTF-8, but TCL source files are interpreted with current system locale, which may not be UTF-8. We must therefore "convert" the entity mapping string to UTF-8 to be certain of TCL's interpretation of its encoding.
Matti Hamalainen <ccr@tnsp.org>
parents:
101
diff
changeset
|
161 |
131
b04ecf8bfb15
urllog: Fix some entity translations.
Matti Hamalainen <ccr@tnsp.org>
parents:
129
diff
changeset
|
162 set urllog_ent_str "-|-|'|'|—|-|‏||—|-|–|--|‪||‬|" |
b04ecf8bfb15
urllog: Fix some entity translations.
Matti Hamalainen <ccr@tnsp.org>
parents:
129
diff
changeset
|
163 append urllog_ent_str "|‎||å|Ã¥|Å|Ã…|é|é|:|:| | " |
133 | 164 append urllog_ent_str "|”|\"|“|\"|«|<<|»|>>|"|\"" |
131
b04ecf8bfb15
urllog: Fix some entity translations.
Matti Hamalainen <ccr@tnsp.org>
parents:
129
diff
changeset
|
165 append urllog_ent_str "|ä|ä|ö|ö|Ä|Ä|Ö|Ö|&|&|<|<|>|>" |
b04ecf8bfb15
urllog: Fix some entity translations.
Matti Hamalainen <ccr@tnsp.org>
parents:
129
diff
changeset
|
166 append urllog_ent_str "|ä|ä|å|ö|—|-|'|'|–|-|"|\"" |
b04ecf8bfb15
urllog: Fix some entity translations.
Matti Hamalainen <ccr@tnsp.org>
parents:
129
diff
changeset
|
167 append urllog_ent_str "|||-|’|'|ü|ü|Ü|Ãœ|•|*|€|€" |
223
606c2a48b2ce
urllog: Add one entity translation.
Matti Hamalainen <ccr@tnsp.org>
parents:
222
diff
changeset
|
168 append urllog_ent_str "|”|\"|‘|'" |
102
5425dc418505
urllog: Entity data is now in UTF-8, but TCL source files are interpreted with current system locale, which may not be UTF-8. We must therefore "convert" the entity mapping string to UTF-8 to be certain of TCL's interpretation of its encoding.
Matti Hamalainen <ccr@tnsp.org>
parents:
101
diff
changeset
|
169 set urllog_html_ent [split [encoding convertfrom "utf-8" $urllog_ent_str] "|"] |
0 | 170 |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
171 ### Require packages |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
172 package require sqlite3 |
0 | 173 package require http |
7
50b52294e93e
urllog: Strip ‏ entities from titles; Some work on SSL/https support.
Matti Hamalainen <ccr@tnsp.org>
parents:
4
diff
changeset
|
174 |
0 | 175 ### Binding initializations |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
176 bind pub - !urlfind urllog_pub_urlfind |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
177 bind msg - !urlfind urllog_msg_urlfind |
0 | 178 bind pubm - *.* urllog_checkmsg |
179 bind topc - *.* urllog_checkmsg | |
180 | |
181 | |
182 ### Initialization messages | |
176
eda776bcb7ed
urllog: Bump copyright, and version.
Matti Hamalainen <ccr@tnsp.org>
parents:
160
diff
changeset
|
183 set urllog_message "$urllog_name v$urllog_version (C) 2000-2014 ccr/TNSP" |
0 | 184 putlog "$urllog_message" |
185 | |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
186 ### HTTP module initialization |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
187 ::http::config -useragent "$urllog_name/$urllog_version" |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
188 if {$http_proxy != 0} { |
28 | 189 ::http::config -proxyhost $http_proxy_host -proxyport $http_proxy_port |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
190 } |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
191 |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
192 if {$http_tls_support != 0} { |
28 | 193 package require tls |
89
77e05ce9e9b8
urllog: Add certdir option setting.
Matti Hamalainen <ccr@tnsp.org>
parents:
87
diff
changeset
|
194 ::http::register https 443 [list ::tls::socket -request 1 -require 1 -cadir $http_tls_cadir] |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
195 } |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
196 |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
197 ### SQLite database initialization |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
198 if {[catch {sqlite3 urldb $urllog_db_file} uerrmsg]} { |
28 | 199 putlog " Could not open SQLite3 database '$urllog_db_file': $uerrmsg" |
200 exit 2 | |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
201 } |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
202 |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
203 |
0 | 204 if {$http_proxy != 0} { |
28 | 205 putlog " (Using proxy $http_proxy_host:$http_proxy_port)" |
0 | 206 } |
207 | |
208 if {$urllog_check != 0} { | |
28 | 209 putlog " (Additional URL validity checks enabled)" |
0 | 210 } |
211 | |
212 if {$urllog_verbose != 0} { | |
28 | 213 putlog " (Verbose mode enabled)" |
0 | 214 } |
215 | |
216 #------------------------------------------------------------------------- | |
217 ### Utility functions | |
218 proc urllog_log {arg} { | |
28 | 219 global urllog_logmsg urllog_name |
0 | 220 |
28 | 221 if {$urllog_logmsg != 0} { |
222 putlog "$urllog_name: $arg" | |
223 } | |
0 | 224 } |
225 | |
226 | |
152 | 227 proc urllog_ctime {utime} { |
28 | 228 if {$utime == "" || $utime == "*"} { |
229 set utime 0 | |
230 } | |
231 return [clock format $utime -format "%d.%m.%Y %H:%M"] | |
0 | 232 } |
233 | |
234 | |
235 proc urllog_isnumber {uarg} { | |
28 | 236 foreach i [split $uarg {}] { |
65
31c8c4f50aa6
urllog: Improve urllog_isnumber function.
Matti Hamalainen <ccr@tnsp.org>
parents:
62
diff
changeset
|
237 if {![string match \[0-9\] $i]} { return 0 } |
28 | 238 } |
65
31c8c4f50aa6
urllog: Improve urllog_isnumber function.
Matti Hamalainen <ccr@tnsp.org>
parents:
62
diff
changeset
|
239 return 1 |
0 | 240 } |
241 | |
242 | |
243 proc urllog_msg {apublic anick achan amsg} { | |
28 | 244 global urllog_preferredmsg |
0 | 245 |
28 | 246 if {$apublic == 1} { |
247 putserv "$urllog_preferredmsg $achan :$amsg" | |
248 } else { | |
249 putserv "$urllog_preferredmsg $anick :$amsg" | |
250 } | |
0 | 251 } |
252 | |
253 | |
254 proc urllog_verb_msg {anick achan amsg} { | |
28 | 255 global urllog_verbose |
0 | 256 |
28 | 257 if {$urllog_verbose != 0} { |
258 urllog_msg 1 $anick $achan $amsg | |
259 } | |
0 | 260 } |
261 | |
262 | |
263 proc urllog_convert_ent {udata} { | |
28 | 264 global urllog_html_ent |
115
5db02af76016
urllog: Improve entity conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
265 return [string map -nocase $urllog_html_ent [string map $urllog_html_ent $udata]] |
0 | 266 } |
267 | |
268 | |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
269 proc urllog_escape { str } { |
28 | 270 return [string map {' ''} $str] |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
271 } |
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
272 |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
273 |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
274 proc urllog_sanitize_encoding {uencoding} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
275 regsub -- "^\[a-z\]\[a-z\]_\[A-Z\]\[A-Z\]\." $uencoding "" uencoding |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
276 set uencoding [string tolower $uencoding] |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
277 regsub -- "^iso-" $uencoding "iso" uencoding |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
278 return $uencoding |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
279 } |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
280 |
121
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
281 proc urllog_clean_title {utitle} { |
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
282 if {[catch {set utitle [encoding convertto "iso8859-15" $utitle]} cerrmsg]} { |
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
283 putlog "Could not convert title encoding: $cerrmsg" |
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
284 } |
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
285 return $utitle |
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
286 } |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
287 |
0 | 288 #------------------------------------------------------------------------- |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
289 set urllog_shorturl_str "ABCDEFGHIJKLNMOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789" |
13
e06d41fb69d5
Begin work on converting urllog.tcl to use an SQLite3 database instead of flat file.
Matti Hamalainen <ccr@tnsp.org>
parents:
8
diff
changeset
|
290 |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
291 proc urllog_get_short {utime} { |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
292 global urllog_shorturl_prefix urllog_shorturl_str |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
293 |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
294 set ulen [string length $urllog_shorturl_str] |
0 | 295 |
28 | 296 set u1 [expr $utime / ($ulen * $ulen)] |
297 set utmp [expr $utime % ($ulen * $ulen)] | |
298 set u2 [expr $utmp / $ulen] | |
299 set u3 [expr $utmp % $ulen] | |
0 | 300 |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
301 return "\[ $urllog_shorturl_prefix[string index $urllog_shorturl_str $u1][string index $urllog_shorturl_str $u2][string index $urllog_shorturl_str $u3] \]" |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
302 } |
0 | 303 } |
304 | |
305 | |
306 #------------------------------------------------------------------------- | |
307 proc urllog_chop_url {url} { | |
28 | 308 global urllog_shorturl_orig |
68 | 309 |
28 | 310 if {[string length $url] > $urllog_shorturl_orig} { |
311 return "[string range $url 0 $urllog_shorturl_orig]..." | |
312 } else { | |
313 return $url | |
314 } | |
0 | 315 } |
316 | |
317 #------------------------------------------------------------------------- | |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
318 proc urllog_exists {urlStr urlNick urlHost urlChan} { |
28 | 319 global urldb urlmsg_alreadyknown urllog_shorturl |
0 | 320 |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
321 set usql "SELECT id AS uid, utime AS utime, url AS uurl, user AS uuser, host AS uhost, chan AS uchan, title AS utitle FROM urls WHERE url='[urllog_escape $urlStr]'" |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
322 urldb eval $usql { |
28 | 323 urllog_log "URL said by $urlNick ($urlStr) already known" |
324 if {$urllog_shorturl != 0} { | |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
325 set qstr "[urllog_get_short $uid] " |
28 | 326 } else { |
327 set qstr "" | |
328 } | |
329 append qstr "($uuser/$uchan@[urllog_ctime $utime])" | |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
330 if {[string length $utitle] > 0} { |
121
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
331 set qstr "$urlmsg_alreadyknown - '[urllog_clean_title $utitle]' $qstr" |
28 | 332 } else { |
333 set qstr "$urlmsg_alreadyknown $qstr" | |
334 } | |
335 urllog_verb_msg $urlNick $urlChan $qstr | |
336 return 0 | |
337 } | |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
338 return 1 |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
339 } |
0 | 340 |
18
1e2232135354
More changes for SQLite support.
Matti Hamalainen <ccr@tnsp.org>
parents:
13
diff
changeset
|
341 |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
342 #------------------------------------------------------------------------- |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
343 proc urllog_addurl {urlStr urlNick urlHost urlChan urlTitle} { |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
344 global urldb urllog_shorturl |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
345 |
93
4e02c0219afe
urllog: Insert NULL into title column when we didn't get a title.
Matti Hamalainen <ccr@tnsp.org>
parents:
92
diff
changeset
|
346 if {$urlTitle == ""} { |
4e02c0219afe
urllog: Insert NULL into title column when we didn't get a title.
Matti Hamalainen <ccr@tnsp.org>
parents:
92
diff
changeset
|
347 set uins "NULL" |
4e02c0219afe
urllog: Insert NULL into title column when we didn't get a title.
Matti Hamalainen <ccr@tnsp.org>
parents:
92
diff
changeset
|
348 } else { |
4e02c0219afe
urllog: Insert NULL into title column when we didn't get a title.
Matti Hamalainen <ccr@tnsp.org>
parents:
92
diff
changeset
|
349 set uins "'[urllog_escape $urlTitle]'" |
4e02c0219afe
urllog: Insert NULL into title column when we didn't get a title.
Matti Hamalainen <ccr@tnsp.org>
parents:
92
diff
changeset
|
350 } |
4e02c0219afe
urllog: Insert NULL into title column when we didn't get a title.
Matti Hamalainen <ccr@tnsp.org>
parents:
92
diff
changeset
|
351 set usql "INSERT INTO urls (utime,url,user,host,chan,title) VALUES ([unixtime], '[urllog_escape $urlStr]', '[urllog_escape $urlNick]', '[urllog_escape $urlHost]', '[urllog_escape $urlChan]', $uins)" |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
352 if {[catch {urldb eval $usql} uerrmsg]} { |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
353 urllog_log "$uerrmsg on SQL:\n$usql" |
28 | 354 return 0 |
355 } | |
82
1bbc79f41a1c
urllog: Rename few variables for clarity.
Matti Hamalainen <ccr@tnsp.org>
parents:
81
diff
changeset
|
356 set uid [urldb last_insert_rowid] |
28 | 357 urllog_log "Added URL ($urlNick@$urlChan): $urlStr" |
0 | 358 |
359 | |
28 | 360 ### Let's say something, to confirm that everything went well. |
361 if {$urllog_shorturl != 0} { | |
82
1bbc79f41a1c
urllog: Rename few variables for clarity.
Matti Hamalainen <ccr@tnsp.org>
parents:
81
diff
changeset
|
362 set qstr "[urllog_get_short $uid] " |
28 | 363 } else { |
364 set qstr "" | |
365 } | |
366 if {[string length $urlTitle] > 0} { | |
121
bec98a9f8695
Convert the title encoding when outputting to channel.
Matti Hamalainen <ccr@tnsp.org>
parents:
120
diff
changeset
|
367 urllog_verb_msg $urlNick $urlChan "'[urllog_clean_title $urlTitle]' ([urllog_chop_url $urlStr]) $qstr" |
28 | 368 } else { |
369 urllog_verb_msg $urlNick $urlChan "[urllog_chop_url $urlStr] $qstr" | |
370 } | |
0 | 371 |
28 | 372 return 1 |
0 | 373 } |
374 | |
375 | |
376 #------------------------------------------------------------------------- | |
377 proc urllog_checkurl {urlStr urlNick urlHost urlChan} { | |
28 | 378 global urllog_tlds urllog_check urlmsg_nosuchhost urlmsg_ioerror |
379 global urlmsg_timeout urlmsg_errorgettingdoc urllog_httprep | |
380 global urllog_shorturl_prefix urllog_shorturl urllog_encoding | |
95
687bdd74dfac
urllog: Check if TLS support is enabled when checking if we can fetch title information via HTTP or SSL/HTTP.
Matti Hamalainen <ccr@tnsp.org>
parents:
93
diff
changeset
|
381 global http_tls_support |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
382 |
96
e5a6c27be365
urllog: Comments and cosmetics.
Matti Hamalainen <ccr@tnsp.org>
parents:
95
diff
changeset
|
383 ### Try to guess the URL protocol component (if it is missing) |
28 | 384 set u_checktld 1 |
385 if {[string match "*www.*" $urlStr] && ![string match "http://*" $urlStr] && ![string match "https://*" $urlStr]} { | |
386 set urlStr "http://$urlStr" | |
387 } elseif {[string match "*ftp.*" $urlStr] && ![string match "ftp://*" $urlStr]} { | |
388 set urlStr "ftp://$urlStr" | |
389 } | |
0 | 390 |
95
687bdd74dfac
urllog: Check if TLS support is enabled when checking if we can fetch title information via HTTP or SSL/HTTP.
Matti Hamalainen <ccr@tnsp.org>
parents:
93
diff
changeset
|
391 ### Handle URLs that have an IPv4-address |
687bdd74dfac
urllog: Check if TLS support is enabled when checking if we can fetch title information via HTTP or SSL/HTTP.
Matti Hamalainen <ccr@tnsp.org>
parents:
93
diff
changeset
|
392 if {[regexp "(\[a-z\]+)://(\[0-9\]{1,3})\\.(\[0-9\]{1,3})\\.(\[0-9\]{1,3})\\.(\[0-9\]{1,3})" $urlStr u_match u_proto ni1 ni2 ni3 ni4]} { |
28 | 393 # Check if the IP is on local network |
92
f6f4595856ff
urllog: Cosmetics. Remove useless parenthesis.
Matti Hamalainen <ccr@tnsp.org>
parents:
91
diff
changeset
|
394 if {$ni1 == 127 || $ni1 == 10 || ($ni1 == 192 && $ni2 == 168) || $ni1 == 0} { |
28 | 395 urllog_log "URL pointing to local or invalid network, ignored ($urlStr)." |
396 return 0 | |
397 } | |
398 # Skip TLD check for URLs with IP address | |
399 set u_checktld 0 | |
400 } | |
0 | 401 |
96
e5a6c27be365
urllog: Comments and cosmetics.
Matti Hamalainen <ccr@tnsp.org>
parents:
95
diff
changeset
|
402 ### Check now if we have an ShortURL here ... |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
403 if {[string match "$urllog_shorturl_prefix*" $urlStr]} { |
98
fbbe7ee40e2f
urllog: Improve one informational / error message.
Matti Hamalainen <ccr@tnsp.org>
parents:
97
diff
changeset
|
404 urllog_log "Ignoring ShortURL from $urlNick: $urlStr" |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
405 set uud "" |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
406 set usql "SELECT id AS uid, url AS uurl, user AS uuser, host AS uhost, chan AS uchan, title AS utitle FROM urls WHERE utime=$uud" |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
407 urldb eval $usql { |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
408 |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
409 } |
28 | 410 return 0 |
411 } | |
0 | 412 |
95
687bdd74dfac
urllog: Check if TLS support is enabled when checking if we can fetch title information via HTTP or SSL/HTTP.
Matti Hamalainen <ccr@tnsp.org>
parents:
93
diff
changeset
|
413 ### Get URL protocol component |
687bdd74dfac
urllog: Check if TLS support is enabled when checking if we can fetch title information via HTTP or SSL/HTTP.
Matti Hamalainen <ccr@tnsp.org>
parents:
93
diff
changeset
|
414 set u_proto "" |
160
e3e156911ab4
urllog: Oops, fix a silly typobug.
Matti Hamalainen <ccr@tnsp.org>
parents:
155
diff
changeset
|
415 regexp "(\[a-z\]+)://" $urlStr u_match u_proto |
95
687bdd74dfac
urllog: Check if TLS support is enabled when checking if we can fetch title information via HTTP or SSL/HTTP.
Matti Hamalainen <ccr@tnsp.org>
parents:
93
diff
changeset
|
416 |
28 | 417 ### Check the PORT (if the ":" is there) |
418 set u_record [split $urlStr "/"] | |
419 set u_hostname [lindex $u_record 2] | |
420 set u_port [lindex [split $u_hostname ":"] end] | |
0 | 421 |
28 | 422 if {![urllog_isnumber $u_port] && $u_port != "" && $u_port != $u_hostname} { |
423 urllog_log "Broken URL from $urlNick: ($urlStr) illegal port $u_port" | |
424 return 0 | |
425 } | |
0 | 426 |
28 | 427 # Default to port 80 (HTTP) |
428 if {![urllog_isnumber $u_port]} { | |
429 set u_port 80 | |
430 } | |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
431 |
28 | 432 ### Is it a http or ftp url? (FIX ME!) |
97
366e68ad94df
urllog: Use u_proto variable to check for if the protocol is supported instead of doing useless additional string checking.
Matti Hamalainen <ccr@tnsp.org>
parents:
96
diff
changeset
|
433 if {$u_proto != "http" && $u_proto != "https" && $u_proto != "ftp"} { |
366e68ad94df
urllog: Use u_proto variable to check for if the protocol is supported instead of doing useless additional string checking.
Matti Hamalainen <ccr@tnsp.org>
parents:
96
diff
changeset
|
434 urllog_log "Broken URL from $urlNick: ($urlStr) UNSUPPORTED protocol class ($u_proto)." |
28 | 435 return 0 |
436 } | |
0 | 437 |
28 | 438 ### Check the Top Level Domain (TLD) validity |
439 if {$u_checktld != 0} { | |
440 set u_sane [lindex [split $u_hostname "."] end] | |
441 set u_tld [lindex [split $u_sane ":"] 0] | |
442 set u_found 0 | |
0 | 443 |
28 | 444 if {[string length $u_tld] == 2} { |
445 # Assume all 2-letter domains to be valid :) | |
446 set u_found 1 | |
447 } else { | |
448 # Check our list of known TLDs | |
449 foreach itld $urllog_tlds { | |
450 if {[string match $itld $u_tld]} { | |
451 set u_found 1 | |
452 } | |
453 } | |
454 } | |
0 | 455 |
28 | 456 if {$u_found == 0} { |
457 urllog_log "Broken URL from $urlNick: ($urlStr) illegal TLD: $u_tld." | |
458 return 0 | |
459 } | |
460 } | |
0 | 461 |
28 | 462 set urlStr [string map $urllog_httprep $urlStr] |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
463 |
91
6f4bfd8e9447
urllog: Reorder code and make it simpler by removing duplicate checks.
Matti Hamalainen <ccr@tnsp.org>
parents:
90
diff
changeset
|
464 ### Does the URL already exist? |
6f4bfd8e9447
urllog: Reorder code and make it simpler by removing duplicate checks.
Matti Hamalainen <ccr@tnsp.org>
parents:
90
diff
changeset
|
465 if {![urllog_exists $urlStr $urlNick $urlHost $urlChan]} { |
6f4bfd8e9447
urllog: Reorder code and make it simpler by removing duplicate checks.
Matti Hamalainen <ccr@tnsp.org>
parents:
90
diff
changeset
|
466 return 1 |
6f4bfd8e9447
urllog: Reorder code and make it simpler by removing duplicate checks.
Matti Hamalainen <ccr@tnsp.org>
parents:
90
diff
changeset
|
467 } |
0 | 468 |
28 | 469 ### Do we perform additional optional checks? |
136
76eefceb2b90
Fix additional checks for HTTPS urls.
Matti Hamalainen <ccr@tnsp.org>
parents:
133
diff
changeset
|
470 if {$urllog_check != 0 && (($http_tls_support != 0 && $u_proto == "https") || $u_proto == "http")} { |
76eefceb2b90
Fix additional checks for HTTPS urls.
Matti Hamalainen <ccr@tnsp.org>
parents:
133
diff
changeset
|
471 # Ok |
76eefceb2b90
Fix additional checks for HTTPS urls.
Matti Hamalainen <ccr@tnsp.org>
parents:
133
diff
changeset
|
472 } else { |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
473 # No optional checks, just add the URL, if it does not exist already |
91
6f4bfd8e9447
urllog: Reorder code and make it simpler by removing duplicate checks.
Matti Hamalainen <ccr@tnsp.org>
parents:
90
diff
changeset
|
474 urllog_addurl $urlStr $urlNick $urlHost $urlChan "" |
28 | 475 return 1 |
476 } | |
7
50b52294e93e
urllog: Strip ‏ entities from titles; Some work on SSL/https support.
Matti Hamalainen <ccr@tnsp.org>
parents:
4
diff
changeset
|
477 |
28 | 478 ### Does the document pointed by the URL exist? |
225
cb86368b8fcd
urllog: Change handling of HTTP requests.
Matti Hamalainen <ccr@tnsp.org>
parents:
224
diff
changeset
|
479 if {[catch {set utoken [::http::geturl $urlStr -timeout 6000 -headers {Accept-Encoding identity}]} uerrmsg]} { |
28 | 480 urllog_verb_msg $urlNick $urlChan "$urlmsg_ioerror ($uerrmsg)" |
481 urllog_log "HTTP request failed: $uerrmsg" | |
482 return 0 | |
483 } | |
0 | 484 |
224
aaf433ab696a
urllog: Improve error messages a bit.
Matti Hamalainen <ccr@tnsp.org>
parents:
223
diff
changeset
|
485 set ustatus [::http::status $utoken] |
aaf433ab696a
urllog: Improve error messages a bit.
Matti Hamalainen <ccr@tnsp.org>
parents:
223
diff
changeset
|
486 if {$ustatus == "timeout"} { |
28 | 487 urllog_verb_msg $urlNick $urlChan "$urlmsg_timeout" |
488 urllog_log "HTTP request timed out ($urlStr)" | |
489 return 0 | |
490 } | |
0 | 491 |
224
aaf433ab696a
urllog: Improve error messages a bit.
Matti Hamalainen <ccr@tnsp.org>
parents:
223
diff
changeset
|
492 if {$ustatus != "ok"} { |
28 | 493 urllog_verb_msg $urlNick $urlChan "$urlmsg_errorgettingdoc ([::http::error $utoken])" |
494 urllog_log "Error in HTTP transaction: [::http::error $utoken] ($urlStr)" | |
495 return 0 | |
496 } | |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
497 |
28 | 498 # Fixme! Handle redirects! |
224
aaf433ab696a
urllog: Improve error messages a bit.
Matti Hamalainen <ccr@tnsp.org>
parents:
223
diff
changeset
|
499 set ustatus [::http::status $utoken] |
aaf433ab696a
urllog: Improve error messages a bit.
Matti Hamalainen <ccr@tnsp.org>
parents:
223
diff
changeset
|
500 set uscode [::http::code $utoken] |
28 | 501 set ucode [::http::ncode $utoken] |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
502 set udata [::http::data $utoken] |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
503 array set umeta [::http::meta $utoken] |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
504 ::http::cleanup $utoken |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
505 |
28 | 506 if {$ucode >= 200 && $ucode <= 309} { |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
507 set uenc_doc "" |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
508 set uenc_http "" |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
509 set uencoding "" |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
510 |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
511 # Get information about specified character encodings |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
512 if {[info exists umeta(Content-Type)] && [regexp -nocase {charset\s*=\s*([a-z0-9._-]+)} $umeta(Content-Type) umatches uenc_http]} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
513 # Found character set encoding information in HTTP headers |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
514 } |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
515 |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
516 if {[regexp -nocase -- "<meta.\*\?content=\"text/html.\*\?charset=(\[^\"\]*)\".\*\?/\?>" $udata umatches uenc_doc]} { |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
517 # Found old style HTML meta tag with character set information |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
518 } elseif {[regexp -nocase -- "<meta.\*\?charset=\"(\[^\"\]*)\".\*\?/\?>" $udata umatches uenc_doc]} { |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
519 # Found HTML5 style meta tag with character set information |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
520 } |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
521 |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
522 # Make sanitized versions of the encoding strings |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
523 set uenc_http2 [urllog_sanitize_encoding $uenc_http] |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
524 set uenc_doc2 [urllog_sanitize_encoding $uenc_doc] |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
525 |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
526 # KLUDGE! |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
527 set uencoding $uenc_http2 |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
528 |
210
52cadf5a12b6
urllog: Disable some debug logging.
Matti Hamalainen <ccr@tnsp.org>
parents:
209
diff
changeset
|
529 # putlog "got charsets : http='$uenc_http', doc='$uenc_doc' / sanitized http='$uenc_http2', doc='$uenc_doc2'" |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
530 |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
531 |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
532 # Check if the document has specified encoding |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
533 if {$uenc_doc != ""} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
534 # Does it differ from what HTTP says? |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
535 if {$uenc_http != "" && $uenc_doc != $uenc_http && $uenc_doc2 != $uenc_http2} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
536 # Yes, we will try reconverting |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
537 set uencoding $uenc_doc2 |
28 | 538 } |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
539 } elseif {$uenc_http == ""} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
540 # If _NO_ known encoding of any kind, assume the default of iso8859-1 |
86
4c2b6482c08c
urllog: Different strategy for charset encoding conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
84
diff
changeset
|
541 set uencoding "iso8859-1" |
4c2b6482c08c
urllog: Different strategy for charset encoding conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
84
diff
changeset
|
542 } |
0 | 543 |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
544 # Get the document title, if any |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
545 set urlTitle "" |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
546 if {[regexp -nocase -- "<title>(.\*\?)</title>" $udata umatches urlTitle]} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
547 # If character set conversion is required, do it now |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
548 if {$uencoding != ""} { |
210
52cadf5a12b6
urllog: Disable some debug logging.
Matti Hamalainen <ccr@tnsp.org>
parents:
209
diff
changeset
|
549 # putlog "conversion requested from $uencoding" |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
550 if {[catch {set urlTitle [encoding convertfrom $uencoding $urlTitle]} cerrmsg]} { |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
551 urllog_log "Error in charset conversion: $cerrmsg" |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
552 } |
28 | 553 } |
150
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
554 |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
555 # putlog "xxx: $uencoding : '$urlTitle'" |
52350ed97775
urllog: Cleanups, rename/move some global variables.
Matti Hamalainen <ccr@tnsp.org>
parents:
136
diff
changeset
|
556 # return 0 |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
557 |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
558 # Convert some HTML entities to plaintext and do some cleanup |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
559 set utmp [urllog_convert_ent $urlTitle] |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
560 regsub -all "\r|\n|\t" $utmp " " utmp |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
561 regsub -all " *" $utmp " " utmp |
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
562 set urlTitle [string trim $utmp] |
28 | 563 } |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
564 |
28 | 565 # Rasiatube hack |
566 if {[string match "*/rasiatube/view*" $urlStr]} { | |
567 set rasia 0 | |
118
e5f2961a6145
urllog: Improve rasiatube URL de-mangling.
Matti Hamalainen <ccr@tnsp.org>
parents:
117
diff
changeset
|
568 if {[regexp -nocase -- "<link rel=\"video_src\"\.\*\?file=(http://\[^&\]+)&" $udata umatches utmp]} { |
e5f2961a6145
urllog: Improve rasiatube URL de-mangling.
Matti Hamalainen <ccr@tnsp.org>
parents:
117
diff
changeset
|
569 regsub -all "\/v\/" $utmp "\/watch\?v=" urlStr |
28 | 570 set rasia 1 |
571 } else { | |
118
e5f2961a6145
urllog: Improve rasiatube URL de-mangling.
Matti Hamalainen <ccr@tnsp.org>
parents:
117
diff
changeset
|
572 if {[regexp -nocase -- "SWFObject.\"(\[^\"\]+)\", *\"flashvideo" $udata umatches utmp]} { |
e5f2961a6145
urllog: Improve rasiatube URL de-mangling.
Matti Hamalainen <ccr@tnsp.org>
parents:
117
diff
changeset
|
573 regsub "http:\/\/www.dailymotion.com\/swf\/" $utmp "http:\/\/www.dailymotion.com\/video\/" urlStr |
28 | 574 set rasia 1 |
575 } | |
576 } | |
577 if {$rasia != 0} { | |
578 urllog_log "RasiaTube mangler: $urlStr" | |
579 urllog_verb_msg $urlNick $urlChan "Korjataan haiseva rasiatube-linkki: $urlStr" | |
580 } | |
581 } | |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
582 |
83
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
583 # Check if the URL already exists, just in case we had some redirects |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
584 if {[urllog_exists $urlStr $urlNick $urlHost $urlChan]} { |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
585 urllog_addurl $urlStr $urlNick $urlHost $urlChan $urlTitle |
f171a9fb7b7b
urllog: Split urllog_add function to urllog_exists for checking whether given URL already exists in the database. Use urllog_exists where appropriate.
Matti Hamalainen <ccr@tnsp.org>
parents:
82
diff
changeset
|
586 } |
28 | 587 return 1 |
588 } else { | |
116
4f3edcf72987
urllog: Improvements in document / HTTP encoding handling and conversion.
Matti Hamalainen <ccr@tnsp.org>
parents:
115
diff
changeset
|
589 urllog_verb_msg $urlNick $urlChan "$urlmsg_errorgettingdoc ($ucode)" |
224
aaf433ab696a
urllog: Improve error messages a bit.
Matti Hamalainen <ccr@tnsp.org>
parents:
223
diff
changeset
|
590 urllog_log "Error fetching document: status=$ustatus, code=$ucode, scode=$uscode, url=$urlStr" |
28 | 591 } |
0 | 592 } |
593 | |
594 | |
595 #------------------------------------------------------------------------- | |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
596 |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
597 |
87 | 598 proc urllog_checkmsg {unick uhost uhand uchan utext} { |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
599 global urllog_log_channels |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
600 |
28 | 601 ### Check the nick |
87 | 602 if {$unick == "*"} { |
28 | 603 urllog_log "urllog_checkmsg: nick was wc, this should not happen." |
604 return 0 | |
605 } | |
0 | 606 |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
607 ### Check the channel |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
608 foreach akey [split $urllog_log_channels] { |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
609 if {[string match $akey $uchan]} { |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
610 ### Do the URL checking |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
611 foreach str [split $utext " "] { |
221
b8bf9d7666b6
urllog: Improve URL / link matching.
Matti Hamalainen <ccr@tnsp.org>
parents:
219
diff
changeset
|
612 if {[regexp "((ftp|http|https)://\[^\[:space:\]\]+|^(www|ftp)\.\[^\[:space:\]\]+)" $str ulink]} { |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
613 urllog_checkurl $str $unick $uhost $uchan |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
614 } |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
615 } |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
616 return 0 |
28 | 617 } |
618 } | |
0 | 619 |
28 | 620 return 0 |
0 | 621 } |
622 | |
623 | |
624 #------------------------------------------------------------------------- | |
625 ### Parse arguments, find and show the results | |
626 proc urllog_find {unick uhand uchan utext upublic} { | |
62
6428b1bcb34b
urllog: Remove some global variable references where they are not used.
Matti Hamalainen <ccr@tnsp.org>
parents:
50
diff
changeset
|
627 global urllog_shorturl urldb |
28 | 628 global urllog_showmax_pub urllog_showmax_priv urlmsg_nomatch |
0 | 629 |
28 | 630 if {$upublic == 0} { |
631 set ulimit 5 | |
632 } else { | |
633 set ulimit 3 | |
634 } | |
19
9cf22053e5da
Repair !urlfind functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
18
diff
changeset
|
635 |
28 | 636 ### Parse the given command |
637 urllog_log "$unick/$uhand searched URL: $utext" | |
0 | 638 |
28 | 639 set ftokens [split $utext " "] |
640 set fpatlist "" | |
641 foreach ftoken $ftokens { | |
642 set fprefix [string range $ftoken 0 0] | |
643 set fpattern [string range $ftoken 1 end] | |
128
0d21b9d1d2b9
urllog: Improve search functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
127
diff
changeset
|
644 set qpattern "'%[urllog_escape $fpattern]%'" |
0 | 645 |
28 | 646 if {$fprefix == "-"} { |
128
0d21b9d1d2b9
urllog: Improve search functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
127
diff
changeset
|
647 lappend fpatlist "(url NOT LIKE $qpattern OR title NOT LIKE $qpattern)" |
28 | 648 } elseif {$fprefix == "%"} { |
128
0d21b9d1d2b9
urllog: Improve search functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
127
diff
changeset
|
649 lappend fpatlist "user LIKE $qpattern" |
28 | 650 } elseif {$fprefix == "@"} { |
651 # foo | |
112
fae3dd7a8b20
urllog: Oops, a typo in variable name. Fixed.
Matti Hamalainen <ccr@tnsp.org>
parents:
111
diff
changeset
|
652 } elseif {$fprefix == "+"} { |
128
0d21b9d1d2b9
urllog: Improve search functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
127
diff
changeset
|
653 lappend fpatlist "(url LIKE $qpattern OR title LIKE $qpattern)" |
28 | 654 } else { |
128
0d21b9d1d2b9
urllog: Improve search functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
127
diff
changeset
|
655 set qpattern "'%[urllog_escape $ftoken]%'" |
0d21b9d1d2b9
urllog: Improve search functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
127
diff
changeset
|
656 lappend fpatlist "(url LIKE $qpattern OR title LIKE $qpattern)" |
28 | 657 } |
658 } | |
19
9cf22053e5da
Repair !urlfind functionality.
Matti Hamalainen <ccr@tnsp.org>
parents:
18
diff
changeset
|
659 |
27
6e381916b016
Some fixes in the query mechanisms of QuoteDB and URLLog.
Matti Hamalainen <ccr@tnsp.org>
parents:
20
diff
changeset
|
660 if {[llength $fpatlist] > 0} { |
6e381916b016
Some fixes in the query mechanisms of QuoteDB and URLLog.
Matti Hamalainen <ccr@tnsp.org>
parents:
20
diff
changeset
|
661 set fquery "WHERE [join $fpatlist " AND "]" |
6e381916b016
Some fixes in the query mechanisms of QuoteDB and URLLog.
Matti Hamalainen <ccr@tnsp.org>
parents:
20
diff
changeset
|
662 } else { |
6e381916b016
Some fixes in the query mechanisms of QuoteDB and URLLog.
Matti Hamalainen <ccr@tnsp.org>
parents:
20
diff
changeset
|
663 set fquery "" |
6e381916b016
Some fixes in the query mechanisms of QuoteDB and URLLog.
Matti Hamalainen <ccr@tnsp.org>
parents:
20
diff
changeset
|
664 } |
68 | 665 |
28 | 666 set iresults 0 |
82
1bbc79f41a1c
urllog: Rename few variables for clarity.
Matti Hamalainen <ccr@tnsp.org>
parents:
81
diff
changeset
|
667 set usql "SELECT id AS uid, utime AS utime, url AS uurl, user AS uuser, host AS uhost FROM urls $fquery ORDER BY utime DESC LIMIT $ulimit" |
68 | 668 urldb eval $usql { |
28 | 669 incr iresults |
670 set shortURL $uurl | |
82
1bbc79f41a1c
urllog: Rename few variables for clarity.
Matti Hamalainen <ccr@tnsp.org>
parents:
81
diff
changeset
|
671 if {$urllog_shorturl != 0 && $uid != ""} { |
1bbc79f41a1c
urllog: Rename few variables for clarity.
Matti Hamalainen <ccr@tnsp.org>
parents:
81
diff
changeset
|
672 set shortURL "$shortURL [urllog_get_short $uid]" |
28 | 673 } |
674 urllog_msg $upublic $unick $uchan "#$iresults: $shortURL ($uuser@[urllog_ctime $utime])" | |
675 } | |
676 | |
677 if {$iresults == 0} { | |
678 # If no URLs were found | |
679 urllog_msg $upublic $unick $uchan $urlmsg_nomatch | |
680 } | |
0 | 681 |
28 | 682 return 0 |
0 | 683 } |
684 | |
685 | |
686 #------------------------------------------------------------------------- | |
687 ### Finding binded functions | |
688 proc urllog_pub_urlfind {unick uhost uhand uchan utext} { | |
219
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
689 global urllog_search_channels |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
690 |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
691 foreach akey [split $urllog_search_channels ";"] { |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
692 if {[string match $akey $uchan]} { |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
693 return [urllog_find $unick $uhand $uchan $utext 1] |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
694 } |
4e09bcc48851
urllog: Add settings for specifying channels where URL logging is active, and where !urlfind functionality works (separately, if so desired.)
Matti Hamalainen <ccr@tnsp.org>
parents:
218
diff
changeset
|
695 } |
28 | 696 return 0 |
0 | 697 } |
698 | |
699 | |
700 proc urllog_msg_urlfind {unick uhost uhand utext} { | |
28 | 701 urllog_find $unick $uhand "" $utext 0 |
702 return 0 | |
3
8003090caa35
Lots of code cleanups, add "fixer" for RasiaTube links (which suck) to point directly to Youtube.
Matti Hamalainen <ccr@tnsp.org>
parents:
0
diff
changeset
|
703 } |
0 | 704 |
705 # end of script |